Qwen 3.8 Flash Q4 3.5t/s on 32GB DDR4 + 8GB GPU
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
I only have 32GB DDR4 and a 8GB GPU. Despite that, I downloaded Qwen 3.8 Flash Q4 quant 93GB in size from unsloth. Guess what, it gets up to 3.5t/s on Debian, freshly compiled. Way better than expected, you can actually chat with that model. This is fucking nuts. edit: I just run llama-cli --model…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-27 21:43 · r/LocalLLaMA
Qwen 3.8 Flash Q4 3.5t/s on 32GB DDR4 + 8GB GPU