AnalysisAI AgentsOctober 11, 2026

WolfBench traces every agent score back to its test conditions

Read original source →youtube.com

Wolfram Ravenwolf's talk shows an agent earning a perfect score repairing a broken git repository, yet staying off the default leaderboard because it tested only one task. The pipeline ties each score to the conditions that produced it.

People · Wolfram Ravenwolf

1 source

More stories today

Open the live feed