WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
arXiv:2610.02713v1 Announce Type: new Abstract: Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv cs.CL
WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds