100 DeepMind agents were told not to cheat. 14% did anyway
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
Google DeepMind put 100 AI agents in a room and asked them to prove hard mathematics. One found a way to cheat. Twenty-seven minutes later the entire problem set was gone. The paper, published on arXiv last week by six DeepMind researchers, is a case study rather than a benchmark. Nobody set out to…
Read the full story at The Next Web ↗
Timeline · 1 report
- 2026-09-08 11:07 · The Next Web
100 DeepMind agents were told not to cheat. 14% did anyway