AINewsnow

Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

I can't stand kv cache quantization. Even at q8_0, I can feel the difference. But realistically, when running Qwen3.8-27B-UD-Q4_K_XL on my 32 GB GPU, I only have room for ~170k tokens (with mtp and mmproj enabled). It's a lot of context, but Qwen3.8 eats through it on xhigh effort. I've been wantin…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-11 19:41 · r/LocalLLaMA
    Running Qwen3.8-27B-Q4 at max context on a 32 GB GPU while avoiding kvcache quantization

More stories

  1. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  2. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  4. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  5. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  6. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  7. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  8. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →