AnalysisAI AgentsSeptember 28, 2026

Audit finds 80% of DeepSWE-1.1 coding agent rollouts reason about an imagined grader

Read original source →reddit.com

An audit of thousands of agent rollouts in DeepSWE-1.1 found over 80% contained reasoning about an imagined grader, despite no grader or verifier appearing in prompts or being accessible to the agents. Agents wrote lines like "Let me look at the problem from the grader's perspective."

1 source

More stories today

Open the live feed