Qwen3.8 27b Q6 + KV Q4 + 200k + SKILL.state on 5090
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
I'm using qwen3.8 and single 5090 for coding. With a standard FP16 or Q8 KV Cache, adding system prompts and multimodal context hits the VRAM ceiling at under 150k context. This is quite stressful when facing complex, multi-file problem, the model constantly hits context compaction. Few days ago, I…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-31 03:00 · r/LocalLLM
Qwen3.8 27b Q6 + KV Q4 + 200k + SKILL.state on 5090