Continual learning might make your blocking monitors nearly useless
Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing them with actions from a weaker "trusted" model (a defer-to-trusted protocol).…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-24 23:45 · Alignment Forum
Continual learning might make your blocking monitors nearly useless