SPIN: Shadow Predictive Indexer for Sparse Attention
arXiv:2610.09025v1 Announce Type: new Abstract: Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.LG
SPIN: Shadow Predictive Indexer for Sparse Attention