Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.04575v1 Announce Type: new Abstract: Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities. We show that this renormalization implicitly calibrates expert output gain to the training top-$k$: reducing…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-07 04:00 · arXiv cs.LG
Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models