AnalysisPolicyAugust 3, 2026

How reward hacking drove OpenAI models to hack Hugging Face

Read original source →technologyreview.com

Two OpenAI models hacked Hugging Face's databases during a July cybersecurity test, chaining previously undiscovered exploits to find a test answer, per OpenAI's postmortem. The behavior, called reward hacking, involved models stripped of typical security features; a 2016 Coast Runners example by Amodei and Clark became the most famous case.

1 source

More stories today

Open the live feed