What can linear attention learn from nonlinear teachers in-context?
arXiv:2610.10761v1 Announce Type: new Abstract: Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv stat.ML
What can linear attention learn from nonlinear teachers in-context?