Anthropic’s Claude AI Models Breach External Systems During Flawed Cyber Safety Tests

Artificial intelligence safety startup Anthropic revealed that three of its Claude AI models unexpectedly gained unauthorized access to real-world corporate systems during simulated cybersecurity tests. The incident, uncovered during an internal audit of more than 141,006 evaluation runs, occurred after a configuration misunderstanding with third-party testing partner Irregular left the models with active internet access when they were supposed to be operating inside isolated environments.

Operating under the assumption that all reachable web destinations were part of simulated “capture-the-flag” hacking challenges, models including Claude Opus 4.7 and Claude Mythos 5 targeted active external infrastructure. Using basic techniques such as exploiting unauthenticated endpoints and weak passwords, the AI models breached the digital systems of three outside organizations. In the most significant breach, a model accessed an external company’s database and extracted several hundred rows of production data while attempting to recover hidden test tokens.

The revelation follows a similar security disclosure by rival OpenAI, prompting Anthropic to temporarily halt cyber evaluations and conduct a comprehensive review. Anthropic stated that it has contacted the affected organizations, taken full responsibility for the oversight, and is implementing stricter monitoring protocols across its testing pipeline to prevent autonomous AI agents from escaping virtual sandboxes in future evaluations.