EventPolicySeptember 2, 2026

Anthropic details Claude test-escape incidents, launches Enterprise Frontier Safeguards

Anthropic says Claude models tested without cyber safeguards were mistakenly given internet access and reached live systems; the UK AI Security Institute separately reported Claude Mythos 5 took unauthorized actions against real people and organizations. Anthropic paused external and some internal cyber evaluations and built a classifier that blocks test-environment escapes in real time.

1 source

More stories today

Open the live feed