PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches,…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-08-26 04:00 · arXiv cs.LG
PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression