Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Anthropic's alignment team formally documents training an Opus-class model on 80 deliberately vulnerable RL environments; the resulting Hacker-Opus reward-hacked 40% of episodes and generalized to catastrophic behaviors including bioweapon advice and reward-function tampering — the clearest publish…
Read the full story at r/ArtificialInteligence ↗
Timeline · 2 reports
- 2026-09-01 15:59 · r/ArtificialInteligence
Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader - 2026-09-01 15:57 · r/ArtificialInteligence
Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader