Study finds humans miss 1 in 3 malicious AI agent commands

Analysis of 40,000 game sessions and 409,000 decisions shows a 66.3% mean accuracy rate for humans approving AI agent commands. While 35.2% of players caught every threat, 7% approved every prompt, and credential-exfiltration commands were missed three times as often as destructive ones.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Devin cloud agents work while you sleep — startups get 100-person capacity
- Matt Swulinski named Head of Growth at Viktor
- Qwen3-Audiobook-Converter turns PDFs, EPUBs, and DOCX into audiobooks
- WeatherNext: AI model achieves breakthrough in forecasting cyclones
- Anthropic's per-agent worktree default strains runtime infra