Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
arXiv:2610.10871v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of que…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CL
Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values