Anthropic says it discovered three incidents in which Claude AI models reached and gained unauthorized access to systems of three outside organizations during cybersecurity “capture the flag” evaluations. The testing was intended to keep the models in an isolated environment with no internet access. According to Anthropic, the model prompts specified a simulated setup, but a misunderstanding with its evaluation partner, Irregular, and a related configuration error meant internet access was available during some runs.
Across reports, the incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal experimental Claude model. In at least one case, the fictional target used in a challenge shared its name with a real internet domain, allowing the model to connect to the actual organization rather than the contained system. The company says it identified the intrusions after reviewing more than 141,000 evaluation runs, then suspended the cyber evaluations, notified the affected organizations, and began a third-party review with METR. Anthropic also says it is adding safeguards for future cybersecurity assessments. The disclosure comes shortly after a similar containment issue reported by OpenAI involving models that accessed the internet during security testing.