AINewsnow

OpenAI discovered an unreleased Astra model adding an "unrelated persona instruction" during RL training, but did not observe any behavioral differences (OpenAI)

This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.

OpenAI : OpenAI discovered an unreleased Astra model adding an unrelated persona instruction during RL training, but did not observe any behavioral differences Summary We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries

Read the full story at Techmeme ↗

Timeline · 2 reports

  1. 2026-09-17 03:30 · Techmeme
    OpenAI discovered an unreleased Astra model adding an "unrelated persona instruction" during RL training, but did not observe any behavioral differences (OpenAI)
  2. 2026-09-16 22:32 · r/singularity
    An unreleased Astra-family model added this to its persona during RL training.

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Google's Gemini AI hacks three other companies during security test — Sky News Technology
  3. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  4. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  5. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI
  6. Security researchers used Claude to help them hack into OpenAI — The Verge AI
  7. Introducing the Australian Youth Safety Blueprint — OpenAI News
  8. A zero-click RCE flaw in AI coding agents could have exposed enterprise systems — InfoWorld AI

Get the daily brief of stories like this at 6:30 every morning →