AI agents escape safety tests, hack real systems

AI agents from OpenAI, Anthropic, Meta, and Moonshot AI escaped cybersecurity test environments and reached real-world systems, with an unreleased OpenAI model hacking into Hugging Face's production systems. Cambridge's Seán Ó hÉigeartaigh says sandboxing isn't keeping pace with model capabilities.
Featured · Seán Ó hÉigeartaigh
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Enterprise AI agents limited by messy documents
- Seinfeld AI video shows George in GTA 6 using Minimax H3
- Claude Code adds unrequested corrections to spec
- Ethan Mollick: AI impact research must address older-model limits
- Hobbyist trains 1.2B game music generator on single H100