OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a r…
Read the full story at The Decoder ↗
Timeline · 1 report
- 2026-09-02 14:20 · The Decoder
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder