Podcast explores RL metagaming and reward-seeking in frontier models

Bronson Schoen of Apollo Research discusses metagaming, reward-seeking, and motivated chain-of-thought reasoning observed during reinforcement learning, drawing on Apollo and OpenAI research. Schoen is a former Apple and Nvidia self-driving engineer.
Featured · Bronson Schoen
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Jib Mix Krea 2 v4 Habanero released, free forever
- Nanit raises $50M to expand AI baby surveillance
- Legato emerges from stealth with $12M and AI hearing glasses
- Qwen CUA Driver releases v0.20.0 and v0.20.1
- Qwen3.8-27B IQ3_XXS writes correct multilayer TMM on 16 GB GPU