Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11.
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
Developer's own thread: https://www.reddit.com/r/LocalLLaMA/s/adp1cGZZe9 Code: https://github.com/Inovello/llama.cpp/tree/flashnext-e06 My hardware: 2x RTX 3090, Intel Ultra 7 270k Plus, 192 GB DDR5@5600 MHz Token generation speed went from 20 t/s to 49 t/s. Prompt processing speed is 140 t/s. Prom…
Read the full story at r/LocalLLaMA ↗
Timeline · 7 reports
- 2026-09-14 04:30 · r/LocalLLaMA
Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra - 2026-09-12 21:08 · r/LocalLLaMA
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo - 2026-09-12 16:52 · r/LocalLLaMA
2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s - 2026-09-12 15:27 · r/LocalLLM
Qwen3.8-Flash-Next on a 48GB M5 Pro MBP: ~20 tok/s in chat, but slow prefill makes agent use unpractical - 2026-09-12 08:18 · r/LocalLLaMA
Qwen3.8 Flash Next llama.cpp config tuning - 2026-09-11 23:11 · r/LocalLLM
Qwen3.8-Flash-Next Q6 on MSI MEG Z790 ACE and 6 consumer GPUs (RTX 3090): ~90 tok/s shallow, ~40 tok/s at 80k context - 2026-09-11 15:00 · r/LocalLLaMA
Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11.