HunterBench benchmark ranks LLMs for autonomous pentesting

HunterBench runs frontier and open LLMs as autonomous pentesters on real infrastructure, scoring coverage and exploitation across two labs (Halcyon and Meridian), each out of 500. Each model runs three times per lab, with results averaged; depth is verified by secret markers.
1 source
Cybersecurity by email
Get an email when there's news on Cybersecurity
No news that day, no email.
More stories today
- Open-source RL training with trl and OpenEnv shared
- NVIDIA invests $3.5B in MediaTek, deepens AI partnership
- Neta team explains why their open-source model generated Anne Hathaway-like images
- OpenAI age-verification error deletes adult's account
- South Korea gives citizens free unlimited domestic AI access