35 B MoE runs under 3 GiB RAM
Learned routing slashes the active memory of a 35‑billion‑parameter Mixture‑of‑Experts model to under 3 GiB, turning laptop‑scale inference from fantasy into practice. By predicting which experts will be needed one token ahead, the engine streams only the relevant weights from SSD and never holds t…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-22 05:00 · DEV Community — Machine Learning
35 B MoE runs under 3 GiB RAM