Attention is a learned weighted average, and the cost is in the square
Originally published at cchinchilla.dev . Part 4 of From code to weights, a 12-part series on ML fundamentals for engineers. Part 3 ended on a square. Two tensors in one decoder block carry T twice, [B, H, T, T] : a T × T square per head, H heads, B sequences. At the widget's defaults they were 89%…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-25 13:00 · DEV Community — Machine Learning
Attention is a learned weighted average, and the cost is in the square