AnalysisAI ModelsJuly 23, 2026

GPT-5.5 scores 10.6% on ActiveVision benchmark, humans 96.1%

GPT-5.5 achieves only 10.6% on the new ActiveVision benchmark, while humans score 96.1%. The failure pattern is notable and cannot be fixed by the model writing its own code.

1 source

More stories today

Open the live feed