AINewsnow

5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug.

Claude Opus 5.5 is the new #1 on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started running this in June 2026. It is also the first model to clear 91 out of 100 on our rubrics…

Read the full story at r/ClaudeAI ↗

Timeline · 1 report

  1. 2026-09-26 20:01 · r/ClaudeAI
    5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug.

More stories

  1. DC appeals court sides with Pentagon on blacklist of Anthropic — The Hill Technology
  2. Question about Wan 3 — r/StableDiffusion
  3. AI system helps lab devices ‘talk’ with each other — streamlining research — Nature — Machine Learning
  4. Deploy and manage coding agents at scale with the Unity Gateway CLI — Databricks Blog
  5. Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents
  6. Wtf just happened ⁉️ — r/GeminiAI
  7. Anthropic's Claude Breaks Physics Record With a Nine-Loop Particle Calculation — AlphaSignal
  8. Made an AR Yu-Gi-Oh prototype with ChatGPT, Astra, CLAD, and Lens Studio — r/ChatGPT

Get the daily brief of stories like this at 6:30 every morning →