LaunchDevelopersJuly 26, 2026
DeepSWE benchmark released with 113 contamination-resistant coding tasks

DeepSWE is a benchmark of 113 software engineering tasks written from scratch to avoid training contamination. Each task is a long-horizon problem from a real open source repo, authored by the repo's maintainer.
Featured · James Shi
1 source
More stories today
- TechCrunch podcast explains panic over Chinese AI model Kimi
- Inside look at the relay market powering LLM token resellers and fraud
- Gebru: Machine learning is rebranded statistical learning
- Reddit user compares Claude Code, OpenCode, and Pi with DeepSeek V4 Flash
- Srinivas criticizes Anthropic, praises Grok and Kimi K