AnalysisAI ModelsJuly 23, 2026
GPT-5.5 scores 10.6% on ActiveVision benchmark, humans 96.1%

GPT-5.5 achieves only 10.6% on the new ActiveVision benchmark, while humans score 96.1%. The failure pattern is notable and cannot be fixed by the model writing its own code.
1 source
More stories today
- Hugging Face Uses Chinese Model in Self-Defense
- Claude replays all players in daily racing game Swervle
- How AI guardrails are impeding the work of offensive cybersecurity researchers
- Fugu-Ultra v1.1 beats Fable 5 by 7.9 points in coding and reasoning
- audio.cpp 0.4 adds Higgs Audio v3 TTS 4B and Fish Audio S2 Pro