New benchmark tests AI agents on large-scale refactoring

A new refactoring-focused benchmark from Shanghai Jiao Tong University, Peking University, and Douyin Group finds the best AI coding agent resolves only 41.2% of tasks, highlighting struggles with large-scale refactoring.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Google DeepMind partners with studios to prototype AI gameplay
- TIME: AI refutes Erdős unit distance conjecture, Fields medalist leaves academia
- Seed: minimal, self-modifying agent harness
- Claude Code skills generate diagrams in Obsidian
- Wazuh AI Analyst enhances SOC workflows