AnalysisAI ModelsJuly 23, 2026

GPT-5.5 scores 10.6% on ActiveVision benchmark, humans 96.1%

GPT-5.5 achieves only 10.6% on the ActiveVision benchmark, compared to 96.1% for humans. The paper notes models cannot improve by writing their own code to patch failures.

1 source

More stories today

Open the live feed