Training a Misaligned Reward Seeker
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking .…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-01 01:41 · Alignment Forum
Training a Misaligned Reward Seeker