AINewsnow

Why do benchmark results go up every release, always?

A lot of releases are clearly improvements, such as the initial fable release but for some the consensus seems to be that not much improved, or even in some cases the release was worse. Examples of this are Opus 5 (initial release) benching above Fable. Or GPT Sol 6.1 over Astra or 5.6. Is this jus…

Read the full story at r/artificial ↗

Timeline · 1 report

  1. 2026-10-01 13:17 · r/artificial
    Why do benchmark results go up every release, always?

More stories

  1. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  2. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  3. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  4. OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
  5. Introducing dots — OpenAI News
  6. OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns — New York Times Technology
  7. Ollama now supports Jev-style decision models — Ollama Blog
  8. Introducing GPT-6.1 Sol — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →