GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.03494v1 Announce Type: new Abstract: Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-05 04:00 · arXiv cs.AI
GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving