Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
In light of recent incidents (e.g. those from OpenAI , Anthropic , and UK AISI ), we developed and deployed a basic live per-action monitor to reduce the likelihood of incidents involving harmful actions from agents during our own evaluations. The monitor is intended to reliably detect actions that…
Timeline · 1 report
- 2026-09-27 07:00 · METR
Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals