Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I can't stand kv cache quantization. Even at q8_0, I can feel the difference. But realistically, when running Qwen3.8-27B-UD-Q4_K_XL on my 32 GB GPU, I only have room for ~170k tokens (with mtp and mmproj enabled). It's a lot of context, but Qwen3.8 eats through it on xhigh effort. I've been wantin…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-11 19:41 · r/LocalLLaMA
Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization