Why I'm scared of RL
Summary: First, I give several different angles on how I feel about reinforcement learning: Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding Recent incidents (huggingface etc) and more mundane form…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-23 12:07 · Alignment Forum
Why I'm scared of RL