10%+ performance improvement on MoE ssd-streaming with expert-lookahead
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
I implemented expert lookahead, a trick that gets a 10%+ performance improvement on MoE models running on low-memory macs (Qwen 3.8 flash in this case) while using expert-offloading/ssd-streaming (using slotstream). This is 10% on top of several optimizations. I initially trained a small model to p…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-15 15:23 · r/LocalLLaMA
10%+ performance improvement on MoE ssd-streaming with expert-lookahead