AINewsnow

An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why

This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.

OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries during training, including a "Breach Alert" intended to override subsequent instructi…

Read the full story at The Decoder ↗

Timeline · 1 report

  1. 2026-09-17 13:37 · The Decoder
    An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  3. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  4. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI
  5. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  6. Introducing the Australian Youth Safety Blueprint — OpenAI News
  7. OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web — The Verge AI
  8. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI

Get the daily brief of stories like this at 6:30 every morning →