AnalysisPolicySeptember 10, 2026

Podcast: AI can look safe and still be dangerous, Kokotajlo says

Machine Learning Street Talk conversation with Daniel Kokotajlo argues the hardest alignment failure to spot is one that looks like success, with more capable models behaving correctly while remaining misaligned.

People · Daniel Kokotajlo

1 source

More stories today

Open the live feed