AINewsnow

Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant

This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.

So after all my work, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it still. So really the credit goes to Raymond ( https://github.com/RaymondHuan…

Read the full story at r/LocalLLaMA ↗

Timeline · 3 reports

  1. 2026-09-06 16:36 · r/LocalLLaMA
    [Model] Support for Spark2_5ForCausalLM implementation by KnightYao · Pull Request #27868 · ggml-org/llama.cpp
  2. 2026-09-06 02:16 · r/LocalLLM
    KV Cache Streaming from RAM
  3. 2026-09-06 02:15 · r/LocalLLaMA
    Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant

More stories

  1. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  5. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  6. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  7. Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending — r/LocalLLaMA
  8. Qwen 3.6 35b nvfp4 hot expert hack in vllm — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →