Understanding and Enhancing Kimi Delta Attention [R]
TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-22 10:34 · r/MachineLearning
Understanding and Enhancing Kimi Delta Attention [R]