AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD
arXiv:2610.06927v1 Announce Type: new Abstract: The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep e…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.LG
AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD