AnalysisAI ModelsJuly 29, 2026
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Two API settings on GPT-5.6 — retaining reasoning and enabling compaction — tripled its ARC-AGI-3 scores and improved efficiency, per OpenAI.
4 sources
How enabling two settings tripled our scores on the ARC-AGI-3 benchmarkopenai.com
Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay:...x.com
Schema Harness Achieves ~99% on Arc‑AGI‑3 Publicschema-harness.github.io
Schema: a harness for llms, with Fable+4.8 or GPT 5.6 Sol, (supposedly) achieves 99% and 95.35% respectively on ARC-AGI-3.reddit.com
OpenAI by email
Get an email when OpenAI ships something
More stories today
- 5.6 sol rolled out on Cerebras for select customers
- Codex user: Work quality 10x'd in 3 months with AI voice mode
- Webflow shares lessons on designing agent-ready MCP APIs
- Voice-only demo builds 3D Iron Man suit with ChatGPT and Claude
- xAI introduces SuperGrok Plus subscription at $100 per month