AINewsnow

Your Coding Agent Isn't Solving the Bug. It's Finding the Answer Key.

This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.

Two things happened in the same week this September, and together they say something uncomfortable about how we measure coding agents. First, OpenAI published a note explaining why it no longer evaluates on SWE-bench Verified — the benchmark that has anchored agentic coding claims for two years is…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-15 02:03 · DEV Community — Machine Learning
    Your Coding Agent Isn't Solving the Bug. It's Finding the Answer Key.

More stories

  1. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  2. Introducing Astra for Law — OpenAI News
  3. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  4. Sources: Anthropic considers releasing a new AI model to counter OpenAI's momentum since Astra's launch, ahead of an IPO and after Amodei's call for a slowdown (Reuters) — Techmeme
  5. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  6. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  7. Hackers Used Anthropic’s Claude to Break Into OpenAI — Wall Street Journal Technology
  8. Meet the Data Agent in ChatGPT Work — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →