Gaussian Equivalence for Multi-Head Self-Attention
arXiv:2610.10033v1 Announce Type: new Abstract: A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scor…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv stat.ML
Gaussian Equivalence for Multi-Head Self-Attention