Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]
In February OpenAI stopped reporting SWE-bench Verified and recommended other labs stop too. Every frontier model they tested could reproduce the human-written reference fix, or verbatim details of the problem statement, for some tasks. Progress had slowed to six points in six months and it wasn't…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-20 14:31 · r/MachineLearning
Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]