Strata: fetch experts before they're needed. Up to +34 % decode headroom measured on 2× 3090, more on smaller-VRAM cards
Quick background: I run Qwen3.8-Flash-Next (a big mixture-of-experts model) on 2x RTX 3090 with Strata ( https://github.com/Niko1221/Strata ). Like every MoE engine on consumer hardware, it can't fit all experts in VRAM. About half of them live in system RAM, and when a token needs one that isn't o…
Read the full story at r/LocalLLM ↗