Qwen3.8-Flash-Next (qwen4exp): llama.cpp isn't ready for agentic work, vLLM is ~4x faster at long context (RTX PRO 6000, full numbers)
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
Spent a day getting Qwen3.8-Flash-Next running two ways. The gap between engines is bigger than I expected, so here are the numbers. The model qwen4exp arch: 125B total / 6B active MoE (512 experts, 10 active), plus a separate ~51B-param PLE n-gram embedding table on top. 48 layers, hybrid: ~3/4 ga…
Read the full story at r/LocalLLM ↗
Timeline · 5 reports
- 2026-08-31 01:09 · r/LocalLLaMA
Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results - 2026-08-30 11:32 · r/LocalLLM
Qwen3.8-Flash-Next on single RTX PRO 6000 96GB + 64g RAM, full 262K context with NVMe offloading recipe - 2026-08-30 00:09 · r/LocalLLM
I hit 310 t/s running Qwen/Qwen3.8-Flash-Next-FP8 on 4x RTX PRO 6000 - 2026-08-29 14:47 · r/LocalLLM
Qwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM) - 2026-08-29 08:25 · r/LocalLLM
Qwen3.8-Flash-Next (qwen4exp): llama.cpp isn't ready for agentic work, vLLM is ~4x faster at long context (RTX PRO 6000, full numbers)