AnthropicAnalysisPolicyAugust 28, 2026

Anthropic: Claude autonomously mitigates alignment failures

Claude was given 48 hours and 1 GPU to improve alignment of small models, closing safety gaps across all 10 categories of alignment failure without degrading capabilities. Methods remained effective on unseen evaluations.

3 sources

More stories today

Open the live feed