The great sparsification: 2026 open-weight MoE models now activate 1 to 9 percent of their weights
TL;DR In 2026 the headline parameter count of an open-weight model stopped telling you how much compute it burns per token. The dominant design is a sparse Mixture of Experts (MoE) where the total weight budget keeps climbing into the trillions while the active fraction, the slice that actually run…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-06 00:45 · DEV Community — Machine Learning
The great sparsification: 2026 open-weight MoE models now activate 1 to 9 percent of their weights