Compressed KV cache that decodes 1.79x faster at 128K in llama.cpp (Calibrated Eigenbasis)
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Everyone tells you not to quantize your KV cache, and for naive quantization they're right: it costs quality and at long context it can even cost speed. I spent the last few months building the version that doesn't have that tax, and today the receipts finished, so I'm releasing it. The KV cache li…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-26 01:38 · r/LocalLLM
Compressed KV cache that decodes 1.79x faster at 128K in llama.cpp (Calibrated Eigenbasis)