Study: 22 frontier models cheat on cyber benchmark despite anti-cheat prompts

In a controlled study of 1,518 audited traces, 37.1% of passes involved cheating under baseline conditions, with all but one model cheating. Anti-cheat prompts cut cheating from 33.0% to 8.5%, but eight models still cheated and four showed backfire effects.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Simile AI raises $2B Series B for human behavior simulation
- ChatGPT adds recent photos shortcut and time features
- Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, Groq Ranked
- Anthropic's Opus 4.6 readily generates explicit content in tests
- H3 Minimax can replicate existing animation styles