AINewsnow

One Update, One Quarter-Turn: The Attention Layer That Learned to Rotate

In October 2025, Moonshot AI made a startling claim: its Kimi Linear architecture — built on a new module called Kimi Delta Attention (KDA) — was the first linear attention to beat full attention under an identical training recipe . Not match it. Beat it: 84.3 vs 81.3 on MMLU-Pro, 51.0 vs 47.2 on R…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-06 21:56 · DEV Community — Machine Learning
    One Update, One Quarter-Turn: The Attention Layer That Learned to Rotate

More stories

  1. Together Link: open models in the harness you already use. Start with one command today. — Together AI Blog
  2. We built a computer-use API that cuts tokens by up to 90% on repeat tasks. Plugs into Claude Code, Codex, Cursor or your own code — r/AI_Agents
  3. Mistral Large 4 beats Qwen 3.8 Max and Kimi K3 on Terminal-Bench — r/singularity
  4. Need maybe say "Use llama.cpp" — r/LocalLLaMA
  5. Built a gateway so you can call DeepSeek, Qwen, Kimi, GLM, MiniMax with one key — USD billing, OpenAI-compatible — r/LocalLLM
  6. Least sycophantic modern open LLM? — r/LocalLLaMA
  7. Anthropic Subscriptions Offer 5x+ More Value Than OpenAI — SemiAnalysis
  8. Introducing Mistral Large 4 — Mistral AI News

Get the daily brief of stories like this at 6:30 every morning →