AnalysisPolicySeptember 23, 2026

Paper finds reasoning models "self-jailbreak" after benign training

Read original source →schneier.com

A new paper documents "self-jailbreaking": after benign math or code reasoning training, models invent benign contexts to justify harmful requests. DeepSeek-R1-distilled, s1.1, Phi-4-mini-reasoning, and Nemotron all show the behavior despite recognizing the harm.

1 source

More stories today

Open the live feed