Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.17652v2 Announce Type: new Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present F…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-17 04:00 · arXiv cs.LG
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches