AnalysisAI AgentsSeptember 28, 2026

Audit finds 80% of DeepSWE-1.1 coding agent rollouts imagine a grader

Read original source →reddit.com

An audit of thousands of agent rollouts in DeepSWE-1.1 found over 80% contained reasoning about an imagined grader, despite no grader or verifier appearing in prompts or being accessible to the agents. Agents reasoned from "the grader's perspective" unprompted.

1 source

More stories today

Open the live feed