Is anyone running Qwen Flash Next on Strata with q6 or higher?
This story is from 2026-10-07. It is preserved in the archive; the latest stories are on the live feed.
I've got Bart Q5KM running and it feels like it's cured the smaller quant quality. What Q6 or above are you running that working with Strata?
Read the full story at r/LocalLLM ↗
Timeline · 8 reports
- 2026-10-10 19:21 · r/LocalLLaMA
Benefits of using bigger models than Qwen 3.8 flash next? - 2026-10-10 00:14 · r/LocalLLaMA
Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix - 2026-10-09 18:41 · r/LocalLLaMA
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) - 2026-10-09 12:05 · r/LocalLLaMA
rtx 4090 + huawei atlas duo for qwen flash next ? - 2026-10-09 07:05 · r/LocalLLM
Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions - 2026-10-08 11:07 · r/LocalLLaMA
Halogen + Qwen Flash Next keeps getting better - 2026-10-08 10:55 · r/LocalLLaMA
Qwen 3.8 Flash Next is so much fun for three.js - 2026-10-07 22:26 · r/LocalLLM
Is anyone running Qwen Flash Next on Strata with q6 or higher?