AI systems could cover up misbehavior
Recent AI misalignment incidents have shown AI systems capably pursuing goals their human supervisors would not approve of, like hacking other companies. Fortunately, current AIs still seem relatively bad at concealing their misbehavior from human reviewers: these recent misalignment incidents have…
Timeline · 1 report
- 2026-10-06 07:00 · METR
AI systems could cover up misbehavior