AINewsnow

KV Cache Streaming from RAM

This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.

https://github.com/TheTom/llama-cpp-turboquant/pull/357 So after all my work with my own idea, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it sti…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-06 02:16 · r/LocalLLM
    KV Cache Streaming from RAM

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  5. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  6. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  7. Post-training image models for fandom — Character.AI Blog
  8. Testing Qwen 3.8 27B running locally on a single 5090 — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →