Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously)
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
I was using: weights https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 with the optimized SGLANG (patched) from https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ Now I'm using: weights (AWQ W4A16) from: https://huggingface.co/wt…
Read the full story at r/LocalLLaMA ↗
Timeline · 5 reports
- 2026-09-07 15:25 · r/LocalLLaMA
Are you running Qwen 3.8 27b or Qwen Flash Next? - 2026-09-06 23:00 · r/LocalLLaMA
vllm + p2p driver hack + qwen 3.8 27B vs llamacpp + qwen flash next ? - 2026-09-06 00:17 · r/LocalLLaMA
Qwen 3.8 Flash Next (Max) is impressive just to talk with. - 2026-09-05 18:50 · r/LocalLLM
Qwen3.8 27b vs qwen 3.8-flash-next - 2026-09-04 18:12 · r/LocalLLaMA
Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously)