Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality
Flash-Next Q4, q8 kv, 262k context 4x P100 on PCIe3 x8 slots (4x16GB=64GB total) 128GB DDR4 at 2133MHz E5-2683 v4 (16c/32t at 2.6GHz boost) On average it's about twice as fast as 27B and five times faster than Flash-Next using a customized llama.cpp just to get it to load at all.
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-04 01:33 · r/LocalLLM
Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality