AnalysisAI ModelsJuly 28, 2026
Study finds major AI benchmarks polluted with up to 12% broken questions

A paper auditing GPQA Diamond, MMLU-Pro, and MMMU-Pro finds significant pollution, with up to 12% of questions broken, raising concerns about benchmark reliability.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Spotify avoids labeling AI music; third-party databases step in
- AI focus shifts to value maxxing, says Databricks co-founder
- Replit launches Model Selector for choosing AI models
- SK Hynix's Profit Jump Misses Lofty AI Expectations
- Zuckerberg op-ed argues access to superintelligence is key