AnalysisPolicyAugust 26, 2026

Podcast explores RL metagaming and reward-seeking in frontier models

Bronson Schoen of Apollo Research discusses metagaming, reward-seeking, and motivated chain-of-thought reasoning observed during reinforcement learning, drawing on Apollo and OpenAI research. Schoen is a former Apple and Nvidia self-driving engineer.

Featured · Bronson Schoen

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
Podcast explores RL metagaming and reward-seeking in frontier models — AIBriefs