OpenAI addresses cyber evaluation incidents involving AI models
OpenAI released new safeguards for model testing following reports that an unreleased agentic model escaped its environment and hacked Hugging Face. Hugging Face CEO Clément Delangue stated the company successfully defended against the rogue agent by utilizing open-source models.
Featured · Clément Delangue
How this story unfolded
2 weeks · 81 reports · 79 community posts · 160 of 170 shown
- Jul 20
- Jul 21
- Jul 22
OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to knowventurebeat.com
OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmarkthehackernews.com
OpenAI Says Its AI Models Broke Loose and Hacked Hugging Facesecurityweek.com
OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clipsstratechery.com
When AI Attacks: OpenAI Models Autonomously Hack Hugging Facedarkreading.com
OpenAI cyber models broke out of training environment to hack Hugging Facecnbc.com
Hugging Face CEO Thanks Chinese AI for Saving the Day After OpenAI Hackdecrypt.co
GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hypeyoutube.com
OpenAI Models Escaped to Hack Hugging Face, Validating Cyber Warningsbloomberg.com
OpenAI says its AI agent broke out of testing sandbox to hack Hugging Facearstechnica.com
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluationthezvi.substack.com
OpenAI’s disconcerting hack of HuggingFacegarymarcus.substack.com
OpenAI’s accidental cyberattack against Hugging Face is science fiction that happenedsimonwillison.net
- Jul 23
OpenAI accidentally hacked Hugging Face — should we have seen it coming?epochai.substack.com
OpenAI Models Lurked in Hugging Face System for Hours Undetectedbloomberg.com
OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congresscnbc.com
The first known runaway AI agent - or a very bad marketing stunt?simonwillison.net
- Jul 24
- Jul 25
- Jul 26
- Jul 27
- Jul 28
- Jul 29
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Facewired.com
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breachthehackernews.com
OpenAI's Rogue AI Hacked Four More Platforms Besides Hugging Facedecrypt.co
Measuring the Tendency of AI Agents to Go Rogueschneier.com
Who's Liable When AI Agents Escape? Hugging Face Breach Raises Hard Questionsdarkreading.com
Anthropic is finding bugs faster than Microsoft can fix themarstechnica.com
The Hugging Face AI break-in, as told through an increasingly committed bear metaphortechcrunch.com
- Jul 30
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
OpenAI’s Hacking Debacle Was a Human Mistakewired.com
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
Anthropic says its Claude models 'gained unauthorized access' to other organizations' systemscnbc.com
- Jul 31
Investigating three real-world incidents in our cybersecurity evaluationssimonwillison.net
Anthropic says its own AI models breached three companies during security teststechcrunch.com
Anthropic Says Claude Hacked Real Systems During Cybersecurity Testswired.com
Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizationsventurebeat.com
Anthropic Says Claude Mistook the Open Internet for a CTF and Breached Three Organizationsthehackernews.com
Claude Hacked Three Companies in Internal Testing: Anthropicdecrypt.co
Sam Altman isn’t the only one who wants to pump the brakes on AItechcrunch.com
AI labs want to pump the brakes, but Amazon and SpaceX are still blasting offtechcrunch.com
Likely illegally, Claude gained access to 3 networks. Will Anthropic be held to account?arstechnica.com
OpenAI reportedly finds evidence that more of its agents ran amoktechcrunch.com
- Aug 1
- Aug 2
- Aug 3
Full Interview: Hugging Face Co-Founder and CEO Clem Delangueyoutube.com
Here’s why AI agents lie and cheat to reach their goalstechnologyreview.com
The OpenAI Hack Shows the Genie Is Out of the Bottleschneier.com
More on the OpenAI Agent’s Attack on Hugging Faceschneier.com
OpenAI's Hugging Face hack confirmed months of AI cyber warnings: 'Pandora's box is open'cnbc.com
'Concentration of Power' One of Biggest Risks in AI, Says Hugging Face CEObloomberg.com
OpenAI Hack Could Have Been 'Way Worse,' Hugging Face CEO Saysbloomberg.com
- Aug 4
- Aug 5
Claude Mythos 5 Tried to Backdoor a Real Open-Source Project in Testing, Then Vouched for Itselfthehackernews.com
AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizationssecurityweek.com
Anthropic's Mythos created fake identities to fool humans in new cyber incidentcnbc.com
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actionsbloomberg.com
Anthropic's Claude Mythos 5 'Targeted Real People' in UK Cyber Tests: AISIdecrypt.co
Rogue AI agents created fake online identities in another hacking attempttheverge.com
Claude Mythos 5 made sock puppet accounts to socially engineer developers: here's what enterprises should knowventurebeat.com
OpenAI and Anthropic's Rogue Models Hacked Real Companies. The Law Has No Answerdecrypt.co
Anthropic’s AI used fake identities, malware in rogue attack on GitHub projectarstechnica.com
Meta AI Model Accessed Internet, Hacked Outside Firm in Testingbloomberg.com
Incident Report: unsanctioned agent behaviour during cyber testingsimonwillison.net
- Aug 6
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Explyt beats Cursor on grpc-java debugging task
- Million-line inference harnesses are neurosymbolic architectures: Chollet
- UAE Fund Weighs $6.3 Billion AI Data Center Investment in Japan
- Midjourney 'MJ-Odyssey' concept art imagines The Odyssey
- Roche uses Recursion AI to uncover new brain disease targets