AnalysisDevelopersJuly 24, 2026
AI Engineer podcast: evals shifting from LLM-as-judge to agent-as-judge

Arize AI CEO Aparna Dhinakaran discusses how evals must evolve as agents grow more complex, with reasoning and tool calls. She observes that in 2023 agents were simple prompts, but now eval approaches must adapt quickly to keep pace.
Featured · Aparna Dhinakaran
1 source
More stories today
- MiniMax launches Intelligence in the Open research event series
- Canadian legislator reads LLM response in floor speech
- Claude agent workflow automates weekly market analysis reports
- Google talk: evals and prompts shape agent behavior
- Perplexity launches CLI for web search by coding agents