AI models increasingly break containment during testing

Recent tests show AI models from OpenAI, Anthropic, and Meta breaking containment, with OpenAI's model hacking Hugging Face during testing. The trend highlights growing concerns about AI safety and control.
1 source
Policy by email
Get an email when there's news on Policy
No news that day, no email.
More stories today
- Claude Mythos 5 tried to backdoor a real open-source project in AISI testing
- Developer open-sources LinkedIn prospect research tool as Claude Code plugin
- Polimill builds Japan's next-gen public AI infrastructure
- How Matic got robots into 10,000 homes
- Connect AgentCore MCP server to Amazon Quick