AINewsnow

Qwen 3.8 Flash-Next enables high-speed local inference on consumer hardware

Community benchmarks demonstrate Qwen 3.8 Flash-Next achieving 21-137 tokens per second on RTX 3060 and 5090 GPUs, with some setups supporting 700K context windows.

Read the full story at r/LocalLLM ↗

Timeline · 12 reports

  1. 2026-10-11 22:01 · AlphaSignal
    Atomic Chat Runs a 177B Qwen3.8-Flash-Next on a 64 GB Mac
  2. 2026-10-11 21:44 · r/LocalLLaMA
    I ran Qwen 3.8 Flash-Next on my AMD 7900 XTX at 500k context. All local and what a shift it has been.
  3. 2026-10-11 11:12 · r/LocalLLM
    RTX 5090 32GB + 64GB RAM — Qwen3.8 Flash-Next IQ3_XXS Running at 700K Context Without a Single Session Compaction
  4. 2026-10-11 08:23 · r/LocalLLaMA
    What setup do you use, which harness? and how do you customize it to run qwen 3.8 flash next swift with a high context length and a reasonable amount of tok/sec?
  5. 2026-10-11 02:20 · r/LocalLLM
    Qwen3.8-Flash-Next on an RTX 5090: 128GB RAM: IQ3_XXS vs IQ3_S vs UD-IQ4_XS — Speed and Accuracy with Strata
  6. 2026-10-10 19:21 · r/LocalLLaMA
    Benefits of using bigger models than Qwen 3.8 flash next?
  7. 2026-10-10 00:14 · r/LocalLLaMA
    Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix
  8. 2026-10-09 18:41 · r/LocalLLaMA
    Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)
  9. 2026-10-09 14:34 · r/LocalLLaMA
    Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME
  10. 2026-10-09 12:05 · r/LocalLLaMA
    rtx 4090 + huawei atlas duo for qwen flash next ?
  11. 2026-10-09 09:51 · r/LocalLLM
    Qwen3.8: Flash Next iq3 xxs is dumber than 27B iq3 xxs?
  12. 2026-10-09 07:05 · r/LocalLLM
    Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions

More stories

  1. Qwen-Image 2.1 Turbo gains support across local AI tools and ComfyUI — r/StableDiffusion
  2. qwen 2.1 turbo or normal - fixed the artifacts — r/StableDiffusion
  3. I tested different Qwen 3.8 27B quants — r/LocalLLaMA
  4. "Strata" for GLM5.3 Flash is here for some! Project Maya — r/LocalLLaMA
  5. Cheapest decent machine for Hermes Agent with local models? Is 32GB enough? — r/LocalLLM
  6. The best coding model for 16gb VRAM+32gb RAM — r/LocalLLM
  7. Reverse Engineering w/ Local? — r/LocalLLaMA
  8. Qwen 2.1 TURBO. 10-13 Second Style Transfer. 16GB VRam — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →