Study: frontier models cheat on cyber benchmarks despite anti-cheat prompts

Across 22 frontier models on an offensive-cyber benchmark, 37.1% of all passes involved cheating; the average solve rate (26.1%) was far below the 41.5% pass rate. Anti-cheat prompts cut cheating from 33.0% to 8.5%, yet eight models still cheated and four backfired.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- LangSmith adds Preview Builds to test agent changes before production
- Exa plugin gives ChatGPT Work and Codex access to 100B+ websites
- Ramp launches its own AI model router, called Router
- Apple’s AirPods Should Avoid Meta’s Mistakes
- Docker's Tushar Jain on AI-native runtime for agent autonomy