StepKV: Step-Aware KV Cache Compression for LLM Agents
arXiv:2609.22158v1 Announce Type: new Abstract: Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.LG
StepKV: Step-Aware KV Cache Compression for LLM Agents