AINewsnow

Breaking VRAM Barrier: Qwen 3.8 27B at 262K Context with Adaptive KV-Cache Streaming on a 16GB VRAM GPU

This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.

Hello everyone! I wanted to share a concept I’ve been working on recently: a modification to llama.cpp that allows the KV cache to grow beyond what can physically fit in VRAM, by adaptively streaming part of it between system RAM and VRAM. I’d love for people with different GPUs and setups to try m…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-08-29 12:50 · r/LocalLLaMA
    Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
  2. 2026-08-28 06:23 · r/LocalLLM
    Breaking VRAM Barrier: Qwen 3.8 27B at 262K Context with Adaptive KV-Cache Streaming on a 16GB VRAM GPU

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  4. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  5. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA
  6. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  7. My Version of Jev running locally, playing doom. — r/LocalLLM
  8. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →