Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM. Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they se…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-10 05:45 · r/LocalLLM
Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM - 2026-10-10 05:41 · r/LocalLLaMA
Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM