Qwen 3.8 Flash-Next enables high-speed local inference on consumer hardware
Community benchmarks demonstrate Qwen 3.8 Flash-Next achieving 21-137 tokens per second on RTX 3060 and 5090 GPUs, with some setups supporting 700K context windows.
Read the full story at r/LocalLLM ↗
Timeline · 12 reports
- 2026-10-11 22:01 · AlphaSignal
Atomic Chat Runs a 177B Qwen3.8-Flash-Next on a 64 GB Mac - 2026-10-11 21:44 · r/LocalLLaMA
I ran Qwen 3.8 Flash-Next on my AMD 7900 XTX at 500k context. All local and what a shift it has been. - 2026-10-11 11:12 · r/LocalLLM
RTX 5090 32GB + 64GB RAM — Qwen3.8 Flash-Next IQ3_XXS Running at 700K Context Without a Single Session Compaction - 2026-10-11 08:23 · r/LocalLLaMA
What setup do you use, which harness? and how do you customize it to run qwen 3.8 flash next swift with a high context length and a reasonable amount of tok/sec? - 2026-10-11 02:20 · r/LocalLLM
Qwen3.8-Flash-Next on an RTX 5090: 128GB RAM: IQ3_XXS vs IQ3_S vs UD-IQ4_XS — Speed and Accuracy with Strata - 2026-10-10 19:21 · r/LocalLLaMA
Benefits of using bigger models than Qwen 3.8 flash next? - 2026-10-10 00:14 · r/LocalLLaMA
Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix - 2026-10-09 18:41 · r/LocalLLaMA
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) - 2026-10-09 14:34 · r/LocalLLaMA
Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME - 2026-10-09 12:05 · r/LocalLLaMA
rtx 4090 + huawei atlas duo for qwen flash next ? - 2026-10-09 09:51 · r/LocalLLM
Qwen3.8: Flash Next iq3 xxs is dumber than 27B iq3 xxs? - 2026-10-09 07:05 · r/LocalLLM
Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions