Qwen3.8-Flash-Next (125B MoE) at 36–45 tok/s and ~1200 t/s prefill on one Ryzen AI Max+ 395, on Windows: open-source runtime + installer
This story is from 2026-09-27. It is preserved in the archive; the latest stories are on the live feed.
I've been building Strix Llama , a patched llama.cpp runtime plus a desktop app for one specific combination: AMD Strix Halo (Ryzen AI Max+ 395, 128 GB) running Qwen3.8-Flash-Next (125B total, ~6B active, Unsloth's UD-IQ4_XS, 94 GB), on Windows . No Linux, no ROCm install, no build tools: one insta…
Read the full story at r/LocalLLM ↗