AINewsnow

Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM

Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4_K_M with llama.cpp at ~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM. Speeds start at ~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they se…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-10 05:45 · r/LocalLLM
    Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM
  2. 2026-10-10 05:41 · r/LocalLLaMA
    Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM

More stories

  1. feat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp — r/LocalLLaMA
  2. 128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build. — r/LocalLLM
  3. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  4. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  5. Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M) — r/LocalLLaMA
  6. Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B? — r/LocalLLaMA
  7. Best uncensored version of Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S ? — r/LocalLLM
  8. China’s open-weight AI models are winning global users. Who is capturing the value? — South China Morning Post Tech

Get the daily brief of stories like this at 6:30 every morning →