AINewsnow

Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open). This is a fir…

Read the full story at r/LocalLLaMA ↗

Timeline · 4 reports

  1. 2026-10-06 11:09 · r/LocalLLM
    Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature
  2. 2026-10-06 11:08 · r/LocalLLaMA
    Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature
  3. 2026-10-05 23:48 · r/LocalLLM
    RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to
  4. 2026-10-05 15:25 · r/LocalLLaMA
    Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

More stories

  1. OpenAI will watermark ChatGPT outputs by default—but only in the EU — Ars Technica AI
  2. Meta and Microsoft take steps to reduce employee usage of Claude AI — Hacker News Front Page
  3. I trust Anthropic with my data. I didn't agree to share it with Meta, TikTok and Google. (IMDEA study) — r/ClaudeAI
  4. Jev for beginners: how to use it and what to build — How I AI
  5. AI Agents Are Moving Into the Real World — The AI Daily Brief
  6. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  7. Introducing Mistral Large 4 — Mistral AI News
  8. Together Link: open models in the harness you already use. Start with one command today. — Together AI Blog

Get the daily brief of stories like this at 6:30 every morning →