AINewsnow

Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time

I maintain MindTrial and tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the xhigh effort label and skip no tasks. Model Passed…

Read the full story at r/ClaudeAI ↗

Timeline · 1 report

  1. 2026-10-04 02:11 · r/ClaudeAI
    Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time

More stories

  1. Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows — AWS Machine Learning Blog
  2. Kimi K3: A Claude clone or something else? — CoreWeave Blog
  3. Implementing Multi-Environment Access for Claude Platform on AWS — AWS Machine Learning Blog
  4. The Museum of Lost Things | Short Film by Claude (Minimax H3) NO user input. — r/ClaudeAI
  5. AI Frontier: Google Gemini-4 Argon — AI Supremacy
  6. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  7. ChatGPT Pro ($100/month) vs Claude Max 5x ($100/month) which is better for coding + technical work? — r/ChatGPTPro
  8. DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness — MarkTechPost

Get the daily brief of stories like this at 6:30 every morning →