AnalysisAI ModelsJuly 28, 2026

Study finds major AI benchmarks polluted with up to 12% broken questions

A paper auditing GPQA Diamond, MMLU-Pro, and MMMU-Pro finds significant pollution, with up to 12% of questions broken, raising concerns about benchmark reliability.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed