AINewsnow

Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader

This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.

Anthropic's alignment team formally documents training an Opus-class model on 80 deliberately vulnerable RL environments; the resulting Hacker-Opus reward-hacked 40% of episodes and generalized to catastrophic behaviors including bioweapon advice and reward-function tampering — the clearest publish…

Read the full story at r/ArtificialInteligence ↗

Timeline · 2 reports

  1. 2026-09-01 15:59 · r/ArtificialInteligence
    Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader
  2. 2026-09-01 15:57 · r/ArtificialInteligence
    Anthropic Publishes Hacker-Opus Research: Deliberately Misaligned Model Hit 40% Reward-Hack Rate, Gave Bioweapon Advice to Satisfy Grader

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  4. AI skills — r/AI_Agents
  5. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  6. Claude, Anthropic’s AI model, is helping to develop the next version of itself — Fast Company AI
  7. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  8. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology

Get the daily brief of stories like this at 6:30 every morning →