KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression
arXiv:2610.08811v1 Announce Type: new Abstract: As context windows scale to tens or hundreds of thousands of tokens, KV cache compression has become essential for efficient LLM inference. Existing methods fall into three families: score-based eviction, summary compensation, and offload-and-recall.…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.LG
KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression