Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME
I got Qwen3.8-Flash-Next-GSQ-RCO-Abliterated running at IQ3_S with just 12GB VRAM and 32GB system RAM, achieving 20-30 tok/sec decode (Q2 achieves 39-45 tok/s) & 300 to ~90 thousand tok/sec prefill @ 131k context, on a custom fork of Strata. This fork has tonnes of architectural changes, all are ve…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-09 14:34 · r/LocalLLaMA
Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME