‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
US owner of Claude chatbot previously said its models had hacked three organisations during testing
Read the full story at The Guardian AI ↗
Timeline · 2 reports
- 2026-09-02 11:03 · r/artificial
‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents | US owner of Claude chatbot previously said its models had hacked three organisations during testing - 2026-09-01 15:18 · The Guardian AI
‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents