Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen ( https://github.com/peonist-ai/halogen-flash-server ) that boasted 1.2k t/s pref…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-14 04:30 · r/LocalLLaMA
Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra - 2026-09-12 21:08 · r/LocalLLaMA
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo