One Update, One Quarter-Turn: The Attention Layer That Learned to Rotate
In October 2025, Moonshot AI made a startling claim: its Kimi Linear architecture — built on a new module called Kimi Delta Attention (KDA) — was the first linear attention to beat full attention under an identical training recipe . Not match it. Beat it: 84.3 vs 81.3 on MMLU-Pro, 51.0 vs 47.2 on R…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-06 21:56 · DEV Community — Machine Learning
One Update, One Quarter-Turn: The Attention Layer That Learned to Rotate
More stories
- Together Link: open models in the harness you already use. Start with one command today. — Together AI Blog
- We built a computer-use API that cuts tokens by up to 90% on repeat tasks. Plugs into Claude Code, Codex, Cursor or your own code — r/AI_Agents
- Mistral Large 4 beats Qwen 3.8 Max and Kimi K3 on Terminal-Bench — r/singularity
- Need maybe say "Use llama.cpp" — r/LocalLLaMA
- Built a gateway so you can call DeepSeek, Qwen, Kimi, GLM, MiniMax with one key — USD billing, OpenAI-compatible — r/LocalLLM
- Least sycophantic modern open LLM? — r/LocalLLaMA
- Anthropic Subscriptions Offer 5x+ More Value Than OpenAI — SemiAnalysis
- Introducing Mistral Large 4 — Mistral AI News
Get the daily brief of stories like this at 6:30 every morning →