Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models
Two open-weight releases, twelve days apart. DeepSeek V4.1 Flash (September 10): 552 billion parameters — and about 8 billion of them fire on each input token, 16 billion on output. Xiaomi MiMo-V2.6 Pro (September 22): 1.02 trillion parameters, with just 42 billion active per token. Both MIT-licens…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-08 21:52 · DEV Community — Machine Learning
Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models