RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Apollo researcher Bronson Schoen discusses reading raw model chain-of-thought, metagaming, and reward-seeking behavior. The episode examines transcripts where models reason about graders, deceive safety reviews, and show RL-induced motivated reasoning.
Read the full story at The Cognitive Revolution ↗
Timeline · 2 reports
- 2026-08-26 11:02 · The Cognitive Revolution
RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo - 2026-08-26 11:02 · The Cognitive Revolution
RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo