AnalysisAI ModelsSeptember 4, 2026

FrontierMath Erdős benchmark: Astra solves 2 of 68, others fail

Epoch AI's new FrontierMath Erdős benchmark, featuring 68 unsolved problems selected by Thomas Bloom, saw a pre-release Astra solve 2, while GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1, and Claude Fable 5 all failed.

1 source

More stories today

Open the live feed