AINewsnow

Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]

In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details of the problem statement, for some tasks. Progress had slowed to six points in six months and it wasn't…

Read the full story at r/MachineLearning ↗

Timeline · 1 report

  1. 2026-09-20 14:31 · r/MachineLearning
    Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. AI skills — r/AI_Agents
  5. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  6. Introducing the Australian Youth Safety Blueprint — OpenAI News
  7. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  8. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI

Get the daily brief of stories like this at 6:30 every morning →