Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF)
Coverage of "Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF)" from 3 sources, with a live timeline of who reported what and when.
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-20 01:47 · r/LocalLLM
I benchmarked 12 quantizations of Qwen3.8-Flash-Next on Strix Halo — here's what actually works - 2026-09-19 21:49 · r/LocalLLaMA
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0) - 2026-09-19 01:42 · r/LocalLLM
Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF)