5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug.
Claude Opus 5.5 is the new #1 on our writing benchmark, and it is not close at all. 2631 Elo. Second one, Fable, is at 2324. That is a 307 point gap, the largest single jump we have recorded since we started running this in June 2026. It is also the first model to clear 91 out of 100 on our rubrics…
Read the full story at r/ClaudeAI ↗
Timeline · 1 report
- 2026-09-26 20:01 · r/ClaudeAI
5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug.