Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)
My journey: 6 tps on UD-Q4_K_XL - hm, this is not right. Claude, find better llama.cpp parameters. 21 tps - that's better, let's see IQ3_XXS 27 tps - nice, but still not my tempo. 51 tps on IQ3_XXS - Niko1221 shares Strata on Reddit. Great! But wait. If there are software gains, there may be more.…
Read the full story at r/LocalLLM ↗
Timeline · 4 reports
- 2026-10-03 01:47 · r/LocalLLM
Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. - 2026-10-01 05:27 · r/LocalLLaMA
Qwen Flash Next MTP work restarted - 2026-09-30 17:15 · r/LocalLLaMA
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks? - 2026-09-30 14:29 · r/LocalLLM
Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really)