AINewsnow

Haiku is 6 points behind Sonnet on SWE-bench and 44 points behind on our internal agent tasks. But on short tasks they tie.

SWE-bench Verified puts Haiku 4.5 6 points behind Sonnet 4.6 (73.3% vs 79.6%). On our agent tasks , same harness , same day: 42.0% vs 85.6%. We bucketed tasks by how many commands Sonnet needed: 1-2 commands: 89.5% vs 89.5% 3-5: 70.7% vs 98.3% 11-20: 46.2% vs 90.8% 41+: 21.2% vs 78.8% Every doublin…

Read the full story at r/ClaudeAI ↗

Timeline · 1 report

  1. 2026-09-22 12:49 · r/ClaudeAI
    Haiku is 6 points behind Sonnet on SWE-bench and 44 points behind on our internal agent tasks. But on short tasks they tie.

More stories

  1. Alibaba Unveils New AI Chip, Outlines Plan for Larger Model — Wall Street Journal Technology
  2. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  3. Amazon blocks Meta’s Muse AI agent — The Verge AI
  4. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  5. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  6. OpenAI forms math advisory group as its AI resolves more than 100 open problems — TechCrunch AI
  7. AIに固有の名前・財布・行動の自由を与えたら、「道具」ではなく「住民」になると思いますか? — r/AI_Agents
  8. British Columbia Sues OpenAI, Alleging ChatGPT Aided Mass School Shooting — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →