AnalysisAI ModelsJuly 23, 2026
GPT-5.5 scores 10.6% on ActiveVision benchmark, humans 96.1%

GPT-5.5 achieves only 10.6% on the ActiveVision benchmark, compared to 96.1% for humans. The paper notes models cannot improve by writing their own code to patch failures.
1 source
More stories today
- ChatGPT Pets get shareable adoption links
- AI Engineer podcast: evals shifting from LLM-as-judge to agent-as-judge
- ChatGPT app renders slowly on iPhone
- OpenAI Build Hour: Valuemaxxing with GPT-5.6
- Enterprises knowingly deployed AI agents without governance controls