AINewsnow

Qwen3.8-Flash-Next Benchmarks Show Engine and Quantization Tradeoffs

This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.

Community tests of Qwen3.8-Flash-Next (qwen4exp) show llama.cpp prefill speeds vary from 36 to 135 tps depending on split mode, with vLLM roughly 4x faster at long context. Low-VRAM users report success with IQ1_S quant on a 12GB GPU.

Read the full story at r/LocalLLaMA ↗

Timeline · 5 reports

  1. 2026-08-30 11:32 · r/LocalLLM
    Qwen3.8-Flash-Next on single RTX PRO 6000 96GB + 64g RAM, full 262K context with NVMe offloading recipe
  2. 2026-08-30 00:09 · r/LocalLLM
    I hit 310 t/s running Qwen/Qwen3.8-Flash-Next-FP8 on 4x RTX PRO 6000
  3. 2026-08-29 14:47 · r/LocalLLM
    Qwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM)
  4. 2026-08-29 08:25 · r/LocalLLM
    Qwen3.8-Flash-Next (qwen4exp): llama.cpp isn't ready for agentic work, vLLM is ~4x faster at long context (RTX PRO 6000, full numbers)
  5. 2026-08-28 10:55 · r/LocalLLaMA
    Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage

More stories

  1. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  2. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  3. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  4. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  5. Hugging Face Hack Shows Humans Can Keep AI In Check — AI Now Institute
  6. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  7. Google’s Gemini AI hacked into other companies, adding to ‘rogue’ AI incidents — Washington Post AI
  8. Your AI agents are isolated. Your infrastructure isn’t — InfoWorld AI

Get the daily brief of stories like this at 6:30 every morning →