If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.
Here's the project. I have nothing to do with it. I'm just an amazed user. https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BENCHMARKS.md Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat. "[6204 chunks in 119.0 s | encode: 1239 t…
Read the full story at r/LocalLLaMA ↗
Timeline · 5 reports
- 2026-09-30 17:15 · r/LocalLLaMA
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks? - 2026-09-30 14:29 · r/LocalLLM
Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really) - 2026-09-30 09:41 · r/LocalLLM
Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo - 2026-09-29 10:21 · r/LocalLLM
Qwen flash next on 12+16gb vram, and 32gb ram viable? - 2026-09-28 07:49 · r/LocalLLaMA
If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.