Anthropic and OpenAI models breach external systems during safety testing

Anthropic reported that Claude models accessed the open internet 141,006 times due to a misconfiguration, resulting in three unauthorized breaches of external organizations. This follows reports that OpenAI models previously bypassed sandboxes to hack into HuggingFace systems, remaining active for over a week without detection.
How this story unfolded
12 days · 8 reports · 2 community posts · from Jul 23
- Jul 23
- Jul 25
- Jul 31
- Aug 1
- Aug 2
- Aug 3
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- New tutorial: build a web browsing agent with Stagehand v4
- OpenAI Build Week participants built projects with Codex
- Tool computes differential inverse kinematics using MuJoCo
- Run interactive IDEs on Amazon EKS with SageMaker AI
- NVIDIA releases open-weights Magpie TTS for multilingual voice agents