Halogen + Qwen Flash Next keeps getting better
This story is from 2026-10-08. It is preserved in the archive; the latest stories are on the live feed.
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai . Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we a…
Read the full story at r/LocalLLaMA ↗
Timeline · 7 reports
- 2026-10-11 02:20 · r/LocalLLM
Qwen3.8-Flash-Next on an RTX 5090: 128GB RAM: IQ3_XXS vs IQ3_S vs UD-IQ4_XS — Speed and Accuracy with Strata - 2026-10-10 19:21 · r/LocalLLaMA
Benefits of using bigger models than Qwen 3.8 flash next? - 2026-10-10 00:14 · r/LocalLLaMA
Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix - 2026-10-09 18:41 · r/LocalLLaMA
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) - 2026-10-09 12:05 · r/LocalLLaMA
rtx 4090 + huawei atlas duo for qwen flash next ? - 2026-10-09 07:05 · r/LocalLLM
Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions - 2026-10-08 11:07 · r/LocalLLaMA
Halogen + Qwen Flash Next keeps getting better