Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080
Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch, and the right cache flags. https://github.com/dtm-beep/qwen38-flash-next-mtp-16gb TLDR: AtomicChat AD-4.27bpw Q4_K_M target + the shared Unsloth MTP…
Read the full story at r/LocalLLaMA ↗
Timeline · 6 reports
- 2026-09-26 13:02 · r/LocalLLM
We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it. - 2026-09-26 08:52 · r/LocalLLaMA
Best current Qwen Flash Next Q4-ish? + worth using? - 2026-09-26 03:42 · r/LocalLLM
PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill - 2026-09-25 06:32 · r/LocalLLaMA
Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some - 2026-09-24 23:28 · r/LocalLLaMA
Is Qwen Flash Next at like Q2 better than 27B at Q4? - 2026-09-23 23:46 · r/LocalLLaMA
Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080