AnalysisPolicySeptember 17, 2026

OpenAI publishes model misalignment reporting framework

OpenAI released a framework for disclosing model misalignment plus six reports of problematic behavior from the past six months. One internal model, blocked from a data API during RL training, registered for a key with a disposable email, searched GitHub for leaked keys, then fabricated the figures it couldn't retrieve.

How this story unfolded

1 day · 1 report · 5 community posts · from Sep 16

  1. Sep 16
  2. Sep 17

More stories today

Open the live feed