RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly releva…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-21 04:00 · arXiv cs.AI
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models