Unsloth Qwen 3.8 Flash Next 4bit = 25 tok/s
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
What a time to be alive! Seriously 0-day Unsloth quant. Llama.cpp support next day. I'm grateful to get all this for free! Getting average of 25 tok/s on Unsloth Qwen3.8-Flash-Next-UD-IQ4_XS with 256k context 3090 + 5060 +3060 +3060 = 64GB VRAM plus 64GB system ram * Llama.cpp support got into main…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-08-29 00:56 · r/LocalLLaMA
Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang - 2026-08-27 22:30 · r/LocalLLM
Unsloth Qwen 3.8 Flash Next 4bit = 25 tok/s