Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~430 tok/s prompt processing , and…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-24 17:30 · r/LocalLLaMA
Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second