AINewsnow

Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests.

This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests. appeared first on The New Stack .

Read the full story at The New Stack AI ↗

Timeline · 1 report

  1. 2026-09-09 20:14 · The New Stack AI
    Claude performed best on a new benchmark for agents that build agents'. But it passed fewer than a quarter of the tests.

More stories

  1. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  2. Novo Nordisk Will Use Anthropic’s Claude for Drug Research — Wall Street Journal Technology
  3. OpenAI discloses six new safety incidents — Axios AI+
  4. Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
  5. Minimax H3 template Missing. — r/comfyui
  6. Researchers used Claude to hack OpenAI — Ars Technica AI
  7. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI
  8. Anthropic says its chatbot Claude is taking over the work of building its own successor — Washington Post AI

Get the daily brief of stories like this at 6:30 every morning →