AnalysisAI ModelsSeptember 15, 2026

Expert re-grading finds broken physics benchmarks, near-saturation

A Yale-led audit of frontier-model physics benchmarks found flawed answer keys; after expert re-grading, scores rose significantly. GPT-5.6 Sol scored 32% on CritPt and 47% on HLE physics, with GPT-6 Astra doing no better.

2 sources

More stories today

Open the live feed