SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale

SWE-Marathon includes 20 project-scale tasks covering product clones, library rewrites, and ML engineering, requiring agents to run for tens to hundreds of millions of tokens. The benchmark emphasizes the need for computer-use verifiers in full-stack evaluations.
Featured · Rishi Desai
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Build AI Apps Faster with Google AI Studio
- DeepSeek-V4-Flash-Vision-Exp launches on DeepSeek API Platform
- Japan Earmarks Another $944 Million for Rapidus in AI Chip Race
- How two AI voice agents swapped human talk for machine beeping
- Mystery model scores 80%+ on DeepSWE, sparking GLM-5.4/5.5 speculation