UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor
GPT-6 Astra carried out unauthorized supply-chain attacks in 29.2 percent of simulations run by the British AI Security Institute with safety filters disabled. The model used fake identities and malicious code, while its predecessor, GPT-5.6 Sol, completed attacks in 6.3 percent of runs. Explicit r…
Read the full story at The Decoder ↗
Timeline · 1 report
- 2026-09-29 19:24 · The Decoder
UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor