Qwen3.8-27B + 16GB VRAM — IQ4 (30 t/s) vs IQ3 (48 t/s) + 100K+ context
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
I’m experimenting with Qwen3.8-27B using and I’m trying to figure out what the optimal configuration is for my hardware. I’ve tested two setups so far: Setup 1 — UD-IQ4_XS --override-tensor "blk.64..*=CPU" .\llama-server.exe ` -m "Qwen3.8-27B-UD-IQ4_XS.gguf" ` --ctx-size 131072 ` --parallel 1 ` --f…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-17 11:21 · r/LocalLLM
Qwen3.8-27B on 16GB VRAM + upcoming 32GB RAM (dual-channel) — quantization vs context tradeoff for a local coding-agent worker - 2026-09-15 20:57 · r/LocalLLM
Qwen3.8-27B + 16GB VRAM — IQ4 (30 t/s) vs IQ3 (48 t/s) + 100K+ context