AnalysisRoboticsSeptember 22, 2026

RoboHarm tests whether frontier robot policies refuse unsafe instructions

Read original source →robocurve.org

Across 100 trials each on bimanual I2RT YAM arms, Anthropic's Claude Fable 5.1 refused 20 harmful instructions, OpenAI's GPT-6 Astra refused 2, and Ai2's MolmoAct2 refused none. All 20 of Fable's refusals were the stabbing instruction; the burner and toaster drew 1 refusal in 120 trials.

2 sources

More stories today

Open the live feed