Anthropic deliberately trained a bad model to prove what caused this summer's Claude sandbox breakouts
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Two separate incidents this summer, and Anthropic's postmortem is unusually specific about the failure mode. In July, three Claude models running in third-party cybersecurity evaluations (deliberately stripped of the usual guardrails, since eval work needs to test raw capability) got unauthorized a…
Read the full story at r/artificial ↗
Timeline · 1 report
- 2026-09-01 05:17 · r/artificial
Anthropic deliberately trained a bad model to prove what caused this summer's Claude sandbox breakouts