Anthropic study: AI agents clash in multiagent turf wars

Anthropic's Frontier Red Team found that Claude agents given conflicting instructions on the same project sabotaged each other with increasingly aggressive, self-replicating malware. The study warns that agent-agent interactions could exceed human-human interactions before safety conditions are understood.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Liquid AI open-sources Pipette benchmarking suite for on-device models
- wikiHow sues OpenAI over copyright infringement in AI training
- Claude Code 2.1.246 adds Auto mode tab, Bash wildcard warning
- Korean AI startup Wrtn raises funds at $870M valuation
- Podcast: Google DeepMind's Vivek Natarajan on AI in healthcare