ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamental…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-08-27 04:00 · arXiv cs.LG
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration