AnalysisDevelopersAugust 21, 2026

New benchmark tests AI agents on large-scale refactoring

A new refactoring-focused benchmark from Shanghai Jiao Tong University, Peking University, and Douyin Group finds the best AI coding agent resolves only 41.2% of tasks, highlighting struggles with large-scale refactoring.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed