Edge0: How a Prerouter and SSD Offloading Let a 35B MoE Model Run in 3 GB of RAM
This story is from 2026-09-18. It is preserved in the archive; the latest stories are on the live feed.
Edge0: How a Prerouter and SSD Offloading Let a 35B MoE Model Run in 3 GB of RAM Running a 35-billion-parameter Mixture-of-Experts (MoE) model on a consumer laptop sounds like a contradiction in terms. At 4-bit precision, those weights alone occupy roughly 19.5 GB — far beyond what most machines ca…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-18 16:04 · DEV Community — Machine Learning
Edge0: How a Prerouter and SSD Offloading Let a 35B MoE Model Run in 3 GB of RAM