AINewsnow

35 B MoE runs under 3 GiB RAM

Learned routing slashes the active memory of a 35‑billion‑parameter Mixture‑of‑Experts model to under 3 GiB, turning laptop‑scale inference from fantasy into practice. By predicting which experts will be needed one token ahead, the engine streams only the relevant weights from SSD and never holds t…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-22 05:00 · DEV Community — Machine Learning
    35 B MoE runs under 3 GiB RAM

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  4. Xiaomi debuts open-weight omnimodal models MiMo-V2.6 Pro and Flash; Pro allegedly performs "on par with Opus 5 and GPT-5.6 Sol across most agent benchmarks" (Xiaomi) — Techmeme
  5. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  6. Lawsuit accuses Anthropic, OpenAI, SpaceXAI, Google of AI pacing 'collusion' — The Hill Technology
  7. Alibaba Unveils AI Chip to Drive Global Data Center Buildout — Bloomberg AI
  8. Grok 4.7 — Hacker News Front Page

Get the daily brief of stories like this at 6:30 every morning →