AnalysisAI ModelsJuly 23, 2026
GPT-5.5 scores 10.6% on ActiveVision benchmark

GPT-5.5 scored only 10.6% on the ActiveVision benchmark, while humans achieved 96.1%. The failure highlights a fundamental limitation that models cannot fix by writing their own code.
1 source
More stories today
- Claude Opus 5 used to build games from scratch in hours
- Robin AI tool reduces dark web research to 30 minutes
- How integrated actuators improve humanoid robot joint performance
- Recursive Superintelligence signs $410 compute deal with Amazon
- Open-source AI financial advisor simulates scenarios