AINewsnow

Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]

Since OpenAI retired SWE-bench Verified in February (every frontier model tested could reproduce reference fixes for some tasks; underspecified tests rewarded knowing the intended fix), I've been trying to write down precisely what a decontamination report can and can't establish. The claim: traini…

Read the full story at r/MachineLearning ↗

Timeline · 1 report

  1. 2026-09-19 17:32 · r/MachineLearning
    Reproduce it, or it doesn't count: why training-side decontamination can't be verified, and what an evaluation-side rule looks like [D]

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. Introducing the Australian Youth Safety Blueprint — OpenAI News
  5. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  6. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  7. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  8. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI

Get the daily brief of stories like this at 6:30 every morning →