When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.04453v1 Announce Type: new Abstract: Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities provide a reliable importance signal. We observe that this assumption brea…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-07 04:00 · arXiv cs.LG
When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models