AINewsnow

We Got 91% With Kimi. Repeatable 91% Is the Hard Part.

This story is from 2026-09-24. It is preserved in the archive; the latest stories are on the live feed.

A peak score tells you what an AI system can do. Production cares about what it can do again tomorrow. We recently got 91% on Terminal-Bench 2.1 using Kimi K3 . Tasks solved: 81/89 Total model cost: $28.72. This is the part where I am apparently supposed to put the number in very large type, add a…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-24 16:29 · DEV Community — AI
    We Got 91% With Kimi. Repeatable 91% Is the Hard Part.

More stories

  1. [R] Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention — r/LocalLLaMA
  2. 2026 AI Model Timeline — r/AI_Agents
  3. Vibe coding Minecraft: January this year vs. today — r/ClaudeAI
  4. Apple just ran a 1T parameter model on four Mac Studios from one wall outlet — r/LocalLLM
  5. DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic — r/LocalLLaMA
  6. Inspur MetaBrain SD200 Ultra Packs 128 Domestic AI Chips for 2.8T Kimi K3 Under 5.85ms/Token — Pandaily
  7. Wow, Mimo 2.6 pro seems to be pretty good, but requires more prompting than Sol — r/LocalLLaMA
  8. Figma Gave GPT-6 Astra a Moonshot. Here's what happened. — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →