Sliding-window beats linear attention
This story is from 2026-08-31. It is preserved in the archive; the latest stories are on the live feed.
Interesting new paper from Alexia Jolicoeur-Martineau (of Tiny Recursive Model fame) and collaborators. They seem to be able to replace quadratic attention with sliding window attention + attention sinks and no post training. This could be big for memory constrained local LLM inference. EDIT: Fixed…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-08-31 16:35 · r/MachineLearning
Sliding-window attention beats linear on long-context reasoning [R] - 2026-08-31 14:12 · r/LocalLLaMA
Sliding-window beats linear attention