How-ToAI AgentsOctober 5, 2026

Talk: evaluating AI agents with LLM judges

Read original source →youtube.com

An LLM correctness judge rejected all 13 reports from a financial analysis agent because the judge graded from its own knowledge while the agent used live web research. Feeding the collected sources to a faithfulness evaluator produced a more accurate grade.

People · Laurie Voss

1 source

More stories today

Open the live feed