Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
a while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3_XXS now runs at ~65 tok/s output and ~ 430 tok/s prompt processing, and…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-27 13:17 · r/LocalLLM
Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second