AINewsnow

Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals

In light of recent incidents (e.g. those from OpenAI , Anthropic , and UK AISI ), we developed and deployed a basic live per-action monitor to reduce the likelihood of incidents involving harmful actions from agents during our own evaluations. The monitor is intended to reliably detect actions that…

Read the full story at METR ↗

Timeline · 1 report

  1. 2026-09-27 07:00 · METR
    Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals

More stories

  1. Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation — The Guardian AI
  2. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  3. Scoop: Top AI companies probing tens of thousands of security incidents — Axios AI+
  4. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  5. Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test? — How I AI
  6. Opus 5.5 — r/ClaudeAI
  7. Anthropic’s new AI system lets lab machines talk to each other and run experiments — Mint AI
  8. Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →