AINewsnow

Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality

Flash-Next Q4, q8 kv, 262k context 4x P100 on PCIe3 x8 slots (4x16GB=64GB total) 128GB DDR4 at 2133MHz E5-2683 v4 (16c/32t at 2.6GHz boost) On average it's about twice as fast as 27B and five times faster than Flash-Next using a customized llama.cpp just to get it to load at all.

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-04 01:33 · r/LocalLLM
    Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality

More stories

  1. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant — r/LocalLLaMA
  5. FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
  6. I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra) — r/LocalLLM
  7. The Rise of Overfit Inference Engines — r/LocalLLaMA
  8. Flash next rig born from mining parts. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →