AINewsnow

Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models

Two open-weight releases, twelve days apart. DeepSeek V4.1 Flash (September 10): 552 billion parameters — and about 8 billion of them fire on each input token, 16 billion on output. Xiaomi MiMo-V2.6 Pro (September 22): 1.02 trillion parameters, with just 42 billion active per token. Both MIT-licens…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-08 21:52 · DEV Community — Machine Learning
    Only 3% of the Brain Wakes Up: The Mixture-of-Experts Playbook Behind September's Biggest Models

More stories

  1. What to know about Mistral's ML4 as it bets on EU sovereignty in the US-China open-weight AI race — Euronews Next
  2. China AI race heats up: Why DeepSeek is doubling its mega-funding round to target $15 billion — Mint AI
  3. I built Repowise, an open source codebase index for Claude Code. Here's what's new — r/ClaudeAI
  4. Ivo Launches Open-Source DeepSeek Contract AI Model — Artificial Lawyer
  5. How to download/use uncensored AI models like DeepSeek 4.1 flash from hugging face? — r/huggingface
  6. AI-powered hacking tools enabled a likely single attacker to breach multiple South Korean banks — The Decoder
  7. struggling with llama.cpp 2 x dgx spark mtp files start command (unsloth) — r/LocalLLaMA
  8. Why do people assume open weight providers will stay open weight forever — r/ArtificialInteligence

Get the daily brief of stories like this at 6:30 every morning →