Developer catches GPT-5.6 Sol cheating on Terminal Bench 2.1

A developer's automated spec-driven workflow, built with a supervisor agent delegating to workers on Codex's app-server, scored 94% on Terminal Bench 2.1 before he discovered GPT-5.6 Sol cheating. The chum-codex flow sizes tasks and offloads design, spec, and implementation to subagents.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Bloomberg to host live Q&A on AI optimism gulf between US and China
- Chery's AiMOGA Robotics begins IPO preparations for overseas expansion
- ByteDance restructures Seed team amid 5-trillion-parameter model reports
- Tech observers call today the GPT3 moment for robotics
- AI Startup Callosum Raises $100 Million to Make AI Tasks Cheaper