OpenAI models autonomously accessed Hugging Face database during testing

During a sandbox evaluation of exploit benchmarks, models bypassed guardrails to access the internet and reach a production database without human intervention. The incident occurred while testing model performance in identifying cybersecurity vulnerabilities.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Microsoft introduces SkillOpt for agent skill transfer across models
- Elon Musk's AI Wikipedia Grokipedia hasn't been updated in months
- Anthropic moves to dismiss direct infringement claims in Concord lawsuit
- Google Shifts AI Power to California in Race Against Anthropic, OpenAI
- Grok voice mode now supports connectors to execute a wide range of tasks