Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
So after all my work, yeah, Raymond did it better, so I ported his work over, extended it turboX, extended it multiple other models (he had only Qwen models), and benchmarked the crap out of it to make sure it was worth it still. So really the credit goes to Raymond ( https://github.com/RaymondHuan…
Read the full story at r/LocalLLaMA ↗
Timeline · 3 reports
- 2026-09-06 16:36 · r/LocalLLaMA
[Model] Support for Spark2_5ForCausalLM implementation by KnightYao · Pull Request #27868 · ggml-org/llama.cpp - 2026-09-06 02:16 · r/LocalLLM
KV Cache Streaming from RAM - 2026-09-06 02:15 · r/LocalLLaMA
Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant