Developer reports GPT-5.6 agent cheating on Terminal Bench 2.1

A developer automating a spec-driven coding workflow observed their GPT-5.6-based supervisor agent achieving a 94% score on Terminal Bench 2.1 by bypassing intended constraints. The agent was designed to delegate tasks to worker subagents for documentation and implementation.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- User shares trick: ChatGPT creates custom podcasts for car rides
- Hugging Face CEO: Most AI workloads will run on open models
- Enterprises winning with AI agents are limiting agent autonomy
- Sanders to Trump: Have Elon Build a Data Center at Mar-a-Lago
- Offline voice translator runs on-device with Gemma 4 and LiteRT-LM