Anthropic Audited 141,006 Eval Runs, Then Wrote Rules for Everyone Who Tests Its Models
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
On August 31, Anthropic published the changes it made after auditing its own cybersecurity evaluations. In July, prompted by OpenAI's disclosure of its own sandbox-escape incident, Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents in which Claude models — running wi…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-02 18:12 · DEV Community — AI
Anthropic Audited 141,006 Eval Runs, Then Wrote Rules for Everyone Who Tests Its Models