Qwen3-Coder 30B on RTX 3080 20GB — KV cache stability and context tuning
I’m running Qwen3-Coder-30B-A3B-Instruct Q3_K_M GGUF in llama.cpp on an RTX 3080 20GB for OpenCode. My stable config so far is: K cache: q8_0 V cache: f16 Flash Attention: off GPU layers: all Parallel: 1 Threads: 8 Continuous batching: on I originally tried: K: q8_0 V: q8_0 Flash Attention: on but…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-24 23:05 · r/LocalLLM
Qwen3-Coder 30B on RTX 3080 20GB — KV cache stability and context tuning