AnalysisAI ModelsJuly 22, 2026

Deep-dive finds AI labs 'pelicanmaxxing' on pelican-bicycle benchmark

Analysis by Dylan Castillo investigates whether AI labs deliberately train models to perform well on the 'pelican riding a bicycle' benchmark, finding signs of targeted optimization. The investigation responds to Simon Willison's informal benchmark and raises questions about benchmark integrity.

Featured · Dylan Castillo

2 sources

More stories today

Open the live feed