AnalysisAI ModelsAugust 21, 2026

New benchmark shows coding agents fail at large-scale refactoring

Read original source →thenewstack.io

A new refactoring-focused benchmark from Shanghai Jiao Tong University, Peking University, and Douyin Group found the best model resolves only 41.2% of tasks. In SWE Refactor Bench, 88 of 520 runs passed all fixed tests, but only 28 survived the full three-stage evaluation.

2 sources

More stories today

Open the live feed