AINewsnow

Qwen3.8-Flash-Next Open Weights Model Draws Early Community Benchmarks

This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.

Qwen released Qwen3.8-Flash-Next, a 125B-token multimodal MoE with 6B active parameters, previewing Qwen4 architecture. Community tests show llama.cpp support merged, with varying speeds and VRAM usage across hardware, while vLLM is reportedly faster at long context.

Read the full story at r/LocalLLM ↗

Timeline · 9 reports

  1. 2026-08-29 14:47 · r/LocalLLM
    Qwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM)
  2. 2026-08-29 13:32 · r/LocalLLM
    Honey, i shrunk Qwen3. 8-Flash-Next
  3. 2026-08-29 08:25 · r/LocalLLM
    Qwen3.8-Flash-Next (qwen4exp): llama.cpp isn't ready for agentic work, vLLM is ~4x faster at long context (RTX PRO 6000, full numbers)
  4. 2026-08-28 10:55 · r/LocalLLaMA
    Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
  5. 2026-08-28 02:25 · r/LocalLLM
    Running Qwen3.8-Flash-Next at Full 262K Context on a 128GB MacBook
  6. 2026-08-27 19:34 · r/LocalLLaMA
    llama.cpp support for Qwen3.8-Flash-Next has been merged
  7. 2026-08-27 12:34 · r/LocalLLaMA
    Qwen3.8-Flash-Next: Time to Update Those Benchmarks
  8. 2026-08-26 23:52 · Simon Willison's Weblog
    Qwen3.8-Flash-Next
  9. 2026-08-26 21:11 · r/LocalLLM
    Qwen3.8 flash next on 4x v100s

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  3. Alibaba releases Qwen-Image-2.1, a 7B open-weight model it says outperforms most closed-source models, with native transparency and up to ten reference images (Qwen) — Techmeme
  4. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  5. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  6. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial
  7. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  8. M2 Mac ultra128gb Qwen flash next — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →