AINewsnow

You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.

I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live in system RAM. I actu…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-16 13:24 · r/LocalLLaMA
    You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  3. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLM
  4. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  5. Qwen Developers on X: "Qwen-Image 2.1 is going open source" — r/StableDiffusion
  6. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  7. Qwen Image 2.1 on Comfy: Coming Soon — r/StableDiffusion
  8. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI

Get the daily brief of stories like this at 6:30 every morning →