Qwen3.8 Flash on 12GB VRAM - 15 tokens/s
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
Achieved steady 15 tokens/second output and 100-120 promp processing per second with 12GB GPU (RTX 5070 SFF) with Qwen3.8-Flash-Next-GSQ-RCO-GGUF at 3 bpw (IQ3_XXS which maches BF16 on AIME25). Around 76GB full gguf - while only 47GB needs to be sharded (loaded into VRAM+RAM) rest is done in SSD. O…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-16 17:41 · r/LocalLLaMA
Qwen3.8 Flash on 12GB VRAM - 15 tokens/s