Studies question reliability of AI safety benchmarks
New papers find chain-of-thought monitoring unreliable in implicit-influence settings and that social pressure breaks majority voting in LLM safety panels. A validity audit of agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) questions their interchangeability.
How this story unfolded
3 days · 3 reports · from Aug 3
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Enterprise AI agents limited by messy documents
- Seinfeld AI video shows George in GTA 6 using Minimax H3
- Claude Code adds unrequested corrections to spec
- Ethan Mollick: AI impact research must address older-model limits
- Hobbyist trains 1.2B game music generator on single H100