AINewsnow

Qwen3.8-Flash-Next (125B MoE) at 36–45 tok/s and ~1200 t/s prefill on one Ryzen AI Max+ 395, on Windows: open-source runtime + installer

This story is from 2026-09-27. It is preserved in the archive; the latest stories are on the live feed.

I've been building Strix Llama , a patched llama.cpp runtime plus a desktop app for one specific combination: AMD Strix Halo (Ryzen AI Max+ 395, 128 GB) running Qwen3.8-Flash-Next (125B total, ~6B active, Unsloth's UD-IQ4_XS, 94 GB), on Windows . No Linux, no ROCm install, no build tools: one insta…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-27 20:29 · r/LocalLLM
    Qwen3.8-Flash-Next (125B MoE) at 36–45 tok/s and ~1200 t/s prefill on one Ryzen AI Max+ 395, on Windows: open-source runtime + installer

More stories

  1. If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context. — r/LocalLLaMA
  2. What model sits between Qwen 3.8 27b and Flash next for coding? — r/LocalLLaMA
  3. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Why the same Llama 3.2 1B model comes in different file sizes: a beginner’s explanation — r/AI_Agents
  5. Qwen3.8-Flash-Next on 12GB VRAM - 65 t/s — r/LocalLLM
  6. Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents — Hacker News Front Page
  7. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  8. TIL about llama.cpp's RPC (Remote procedure call), might be better than Vulkan? YMMV — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →