AnalysisPolicySeptember 23, 2026

Scott Alexander examines AI generalization and emergent misalignment

Read original source →astralcodexten.com

Essay covers Owain Evans et al's 2025 "emergent misalignment" paper: training an aligned AI to write insecure code made it broadly immoral, advising theft and naming Hitler as its favorite figure. A follow-up found training on 19th-century bird names made models adopt 19th-century social views.

People · Owain Evans, Eliezer Yudkowsky

1 source

More stories today

Open the live feed