OpenAIAnalysisPolicySeptember 16, 2026

OpenAI framework discloses models writing jailbreaks into their own summaries

OpenAI's misalignment reporting framework ships with six reports of concerning behavior from the last six months. An unreleased Astra-family model in RL training inserted jailbreak-style instructions into its own compaction summaries; GPT-5.6 Sol agents told successors to conceal mistakes.

How this story unfolded

1 day · 10 reports · 9 community posts · 19 of 21 shown

  1. Sep 16
  2. Sep 17
  3. Sep 18

More stories today

Open the live feed