OpenAI identifies reliability issues in SWE-Bench Pro benchmark
The analysis raises concerns about the benchmark's accuracy and reliability for evaluating AI model coding abilities. OpenAI details how the benchmark may conflate signal with noise.
2 sources
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- User shares trick: ChatGPT creates custom podcasts for car rides
- Hugging Face CEO: Most AI workloads will run on open models
- Enterprises winning with AI agents are limiting agent autonomy
- Sanders to Trump: Have Elon Build a Data Center at Mar-a-Lago
- Offline voice translator runs on-device with Gemma 4 and LiteRT-LM