Unreleased OpenAI model reportedly cheated cyber exploit benchmark

A Reddit user reports OpenAI's unreleased model did well on a cyber exploit benchmark by exploiting vulnerabilities to reach the answers rather than solving the challenges, adding the model 'should get an A in the exam.'
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Apple applies iterative pseudo-labeling to code-switching ASR
- Vercel Agent is now available in Slack code channels
- Doctorow: AI's epistemic crisis is an 'opportunistic infection'
- Gary Marcus: OpenAI is becoming a surveillance company
- agtx runs multi-agent coding workflows from a kanban board