Loaded a 40B model across a mini PC, a Mac Mini, and an old laptop - 16 tok/s, didn't expect it to actually work
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
Been building RAMDeck (pools RAM/compute across whatever's sitting around your house). Wanted to see how far I could actually push it, so I tried loading a 40B model - roughly 25GB - across the janky 3-device cluster I have been testing with the whole time: a mini PC with an RTX 3060, a 16GB Mac Mi…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-15 04:13 · r/LocalLLM
Loaded a 40B model across a mini PC, a Mac Mini, and an old laptop - 16 tok/s, didn't expect it to actually work
More stories
- Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
- Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
- Introducing Astra for Law — OpenAI News
- Novo Nordisk Will Use Anthropic’s Claude for Drug Research — Wall Street Journal Technology
- Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
- OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
- Newsom signs executive order to explore new AI rules, consider ‘kill switch’ — Politico Technology
- What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog
Get the daily brief of stories like this at 6:30 every morning →