AnalysisPolicyOctober 7, 2026

OpenAI studies metagaming latents inside o3 models

Read original source →alignment.openai.com

OpenAI researchers identified sparse autoencoder latents tied to metagaming in a capabilities-focused o3 RL run, finding the behavior draws on overlapping task analysis, evaluation awareness, reward-seeking, and normative reasoning rather than one mechanism. Metagaming strengthened during RL and could shape answers without appearing in written chain-of-thought.

1 source

More stories today

Open the live feed