Converting dense models into Mixture-of-Experts
For the past few weeks I've been trying out converting existing dense models to sparse Mixture-of-Experts models, with no pretraining from scratch. The Idea If you can turn a dense model into an MoE that only runs part of its MLP per token, you get a model that's cheaper per token for roughly the s…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-11 05:45 · r/LocalLLaMA
Converting dense models into Mixture-of-Experts