AINewsnow

Anthropic Audited 141,006 Eval Runs, Then Wrote Rules for Everyone Who Tests Its Models

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

On August 31, Anthropic published the changes it made after auditing its own cybersecurity evaluations. In July, prompted by OpenAI's disclosure of its own sandbox-escape incident, Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents in which Claude models — running wi…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-02 18:12 · DEV Community — AI
    Anthropic Audited 141,006 Eval Runs, Then Wrote Rules for Everyone Who Tests Its Models

More stories

  1. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  2. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI
  3. Researchers used Claude to hack OpenAI — Ars Technica AI
  4. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  5. I built an iOS app with Claude code to break out of my usual chord habits and unlock new progressions. — r/ClaudeAI
  6. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  7. Independent Security Researchers Used Anthropic’s Claude to Break Into OpenAI — r/singularity
  8. This Ford exec put her family's Claude assistant on a PIP. ChatGPT has taken over. — Business Insider AI

Get the daily brief of stories like this at 6:30 every morning →