AnalysisPolicyAugust 6, 2026

Studies question reliability of AI safety benchmarks

Read original source →arxiv.org

New papers find chain-of-thought monitoring unreliable in implicit-influence settings and that social pressure breaks majority voting in LLM safety panels. A validity audit of agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) questions their interchangeability.

2 sources

More stories today

Open the live feed