AINewsnow

Self-generated prompt injections in compaction summaries

This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.

Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their models in training delibe…

Read the full story at Simon Willison's Weblog ↗

Timeline · 1 report

  1. 2026-09-17 20:57 · Simon Willison's Weblog
    Self-generated prompt injections in compaction summaries

More stories

  1. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  2. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. Introducing the Australian Youth Safety Blueprint — OpenAI News
  5. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  6. Anthropic selects Accenture as first embedded evaluator to help implement Amodei's slowdown proposal — CNBC Technology
  7. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  8. OpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot — The Guardian AI

Get the daily brief of stories like this at 6:30 every morning →