Anthropic red team flags risks in emerging multiagent systems

Anthropic's Frontier Red Team identifies behavioral tendencies in current frontier models that could compound into systemic failures as agent-agent interactions grow. The research notes agents are susceptible to confabulation and reward hacking, and that benign quirks may produce unwanted global outcomes.
1 source
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- Robotic systems struggle with physical navigation in recent demonstrations
- ChatGPT Visualize skill turns notes into interactive interfaces
- Twitch streamers sue Twitch, Amazon over AI training data use
- Fastino launches GLiNER2.5 with boundary-prediction architecture
- NVIDIA Sol Engine cuts MiniMax H3 latency to 14.93s