AnalysisAI ModelsJuly 23, 2026

GPT-5.5 scores 10.6% on ActiveVision benchmark, humans 96.1%

GPT-5.5 scored just 10.6% on the new ActiveVision benchmark, while humans achieved 96.1%. The paper highlights that models cannot patch failures by generating their own code, revealing a key limitation.

1 source

More stories today

Open the live feed