Five new papers expose flaws in AI benchmark and safety scoring
An audit of safety benchmarks R-Judge, InjecAgent, AgentHarm and AgentDojo finds their scores quoted interchangeably despite measuring different behaviors. Another paper finds contamination checks uninformative: four flagship models fail them on unmemorizable questions. A third coins "evaluation blindness," silent measurement failures from training to deployment.
How this story unfolded
3 days · 5 reports · from Aug 3
- Aug 3
- Aug 4
- Aug 5
- Aug 6
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Paper examines the limitations of current AI evaluation methods
- OnlyHuman filter list removes AI-generated SEO spam from search results
- Qwen tokenizes 330-line code into 1,609 tokens; Gemma needs 4,258
- LifeOS: open-source AI harness for personal growth and work
- MINIMAX video drops Indiana Jones into Mortal Kombat