AnalysisAI ModelsAugust 6, 2026

Five new papers expose flaws in AI benchmark and safety scoring

An audit of safety benchmarks R-Judge, InjecAgent, AgentHarm and AgentDojo finds their scores quoted interchangeably despite measuring different behaviors. Another paper finds contamination checks uninformative: four flagship models fail them on unmemorizable questions. A third coins "evaluation blindness," silent measurement failures from training to deployment.

How this story unfolded

3 days · 5 reports · from Aug 3

  1. Aug 3
  2. Aug 4
  3. Aug 5
  4. Aug 6

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed