Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)
Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0: decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter) decode at 258,794: 42.9 to 45.0 prefil…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-20 01:47 · r/LocalLLM
I benchmarked 12 quantizations of Qwen3.8-Flash-Next on Strix Halo — here's what actually works - 2026-09-19 21:49 · r/LocalLLaMA
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)