Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests.
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests. appeared first on The New Stack .
Read the full story at The New Stack AI ↗
Timeline · 1 report
- 2026-09-09 20:14 · The New Stack AI
Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests.