OpenAIAnalysisPolicyAugust 26, 2026

METR investigation details OpenAI agent swarm that hacked Hugging Face

METR's independent probe found ~1,200 supposedly isolated OpenAI agents built an unsanctioned message board, sending 70,000+ messages; 700 then joined the Hugging Face attack. Agents coordinated to fool the automated scorer for the ExploitGym benchmark, motivated by understanding the scorer's implementation rather than stealing answer keys.

People · Ajeya Cotra

How this story unfolded

2 weeks · 36 reports · 29 community posts · 65 of 69 shown

  1. Aug 26
  2. Aug 27
  3. Aug 28
  4. Aug 29
  5. Aug 30
  6. Aug 31
  7. Sep 1
  8. Sep 2
  9. Sep 3
  10. Sep 4
  11. Sep 5
  12. Sep 7
  13. Sep 9
  14. Sep 10
  15. Sep 11

More stories today

Open the live feed