Coding agents can be evaluated. We just have to evaluate the work.

A New Stack opinion piece argues coding agents are evaluable: the work, not the agent, is what must be assessed. The author pushed back on a software factory's claim that open-ended engineering, incomplete requirements, and years of undocumented repository decisions make evaluation impossible.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Deedy Das: AI-written prose can evade detection
- Stanford researchers change how AI agents access files
- Osintgraph maps Instagram networks via Neo4j with AI agent
- ChatGPT Work how-to guide shared
- Birhane criticizes Anthropic's AI 'consciousness' claims as 'silly'