Haiku is 6 points behind Sonnet on SWE-bench and 44 points behind on our internal agent tasks. But on short tasks they tie.
SWE-bench Verified puts Haiku 4.5 6 points behind Sonnet 4.6 (73.3% vs 79.6%). On our agent tasks , same harness , same day: 42.0% vs 85.6%. We bucketed tasks by how many commands Sonnet needed: 1-2 commands: 89.5% vs 89.5% 3-5: 70.7% vs 98.3% 11-20: 46.2% vs 90.8% 41+: 21.2% vs 78.8% Every doublin…
Read the full story at r/ClaudeAI ↗