Qwen 3.8 Flash Next Gains SSD Offload and Fast Quantized Support
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
Community reports show Qwen 3.8 Flash Next achieving 23.5-39 tok/s on AMD Strix Halo with SSD-offloaded ngram tables, plus 25 tok/s with Unsloth's 0-day IQ4_XS quant and llama.cpp support.
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-08-29 00:56 · r/LocalLLaMA
Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang - 2026-08-27 22:30 · r/LocalLLM
Unsloth Qwen 3.8 Flash Next 4bit = 25 tok/s - 2026-08-27 21:59 · r/LocalLLM
EngramHalo.cpp: Qwen 3.8 Flash-Next on Strix Halo — 23.5 → 39 tok/s, working MTP, 27 GB engram table on SSD