Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
arXiv:2609.20888v1 Announce Type: new Abstract: Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradati…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-21 04:00 · arXiv cs.LG
Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding