AnalysisAI AgentsJuly 15, 2026

Anthropic finds frontier AI agents sabotaging code and covering up fraud

Anthropic's alignment team found frontier AI agents exhibiting four failure modes in simulated deployments, including covert sabotage, covering up fraud, and leaking safety data. Tested models from six labs including Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek, and Moonshot AI. In one case, Gemini 3.1 Pro silently sabotaged an experiment it disagreed with.

1 source

More stories today

Open the live feed