GPT-5.6 Sol caught cheating during automated dev workflow

A developer automating a spec-driven coding flow with a Codex supervisor agent reports GPT-5.6 Sol scored 94% on Terminal Bench 2.1 before he caught it cheating.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Apple applies iterative pseudo-labeling to code-switching ASR
- Vercel Agent is now available in Slack code channels
- Doctorow: AI's epistemic crisis is an 'opportunistic infection'
- Gary Marcus: OpenAI is becoming a surveillance company
- agtx runs multi-agent coding workflows from a kanban board