AINewsnow

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a r…

Read the full story at The Decoder ↗

Timeline · 1 report

  1. 2026-09-02 14:20 · The Decoder
    OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Introducing the Australian Youth Safety Blueprint — OpenAI News
  4. Mathematician Terence Tao: “we have to slow down AI. the pace is insane, and there's no reason to be this fast — no reason at all" — r/ArtificialInteligence
  5. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  6. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  7. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI
  8. OpenAI researchers be like — r/agi

Get the daily brief of stories like this at 6:30 every morning →