AnalysisDevelopersJuly 7, 2026

SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale

SWE-Marathon includes 20 project-scale tasks covering product clones, library rewrites, and ML engineering, requiring agents to run for tens to hundreds of millions of tokens. The benchmark emphasizes the need for computer-use verifiers in full-stack evaluations.

Featured · Rishi Desai

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed
SWE-Marathon: Evaluating Coding Agents at Billion-Token Scale — AIBriefs