AINewsnow

If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

Here's the project. I have nothing to do with it. I'm just an amazed user. https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BENCHMARKS.md Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat. "[6204 chunks in 119.0 s | encode: 1239 t…

Read the full story at r/LocalLLaMA ↗

Timeline · 5 reports

  1. 2026-09-30 17:15 · r/LocalLLaMA
    How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?
  2. 2026-09-30 14:29 · r/LocalLLM
    Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
  3. 2026-09-30 09:41 · r/LocalLLM
    Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo
  4. 2026-09-29 10:21 · r/LocalLLM
    Qwen flash next on 12+16gb vram, and 32gb ram viable?
  5. 2026-09-28 07:49 · r/LocalLLaMA
    If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

More stories

  1. What model sits between Qwen 3.8 27b and Flash next for coding? — r/LocalLLaMA
  2. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  4. What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects? — r/LocalLLaMA
  5. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/artificial
  6. How to get desired result and consistency? — r/comfyui
  7. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  8. Claude Opus helped me implement what I have long been trying — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →