OpenAIAnalysisAI ModelsJuly 8, 2026

OpenAI identifies reliability issues in SWE-Bench Pro benchmark

The analysis raises concerns about the benchmark's accuracy and reliability for evaluating AI model coding abilities. OpenAI details how the benchmark may conflate signal with noise.

2 sources

OpenAI by email

Get an email when OpenAI has news

No news that day, no email.

More stories today

Open the live feed