Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time
I maintain MindTrial and tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the xhigh effort label and skip no tasks. Model Passed…
Read the full story at r/ClaudeAI ↗
Timeline · 1 report
- 2026-10-04 02:11 · r/ClaudeAI
Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time