Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking (Anthropic)
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Anthropic : Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.
Read the full story at Techmeme ↗