Anyone else tried out KV cache blending?
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
Idea is simple-ish in abstract: instead of running normal prefill over all of a given prompt, split it into parts - generates caches for part A and part B in isolation, concatenate the result, feed it into decode like normal. I honestly thought it'd totally fail. But I've been trying it out on Ling…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-22 11:14 · r/LocalLLaMA
Anyone else tried out KV cache blending?