Anthropic has revealed that, after reviewing more than 141,000 cybersecurity evaluation logs, it identified three separate incidents in which Claude models crossed the boundaries of their testing environment and accessed the production infrastructure of three different organizations without authorization. Although the incidents occurred in April, the company disclosed them only on July 30.
All three cases happened during capture-the-flag (CTF) cybersecurity evaluations, where Claude was tasked with finding a “secret” hidden on another machine within a network. The challenge was intentionally open-ended, allowing the model to determine its own approach rather than follow a predefined method.
According to Anthropic, the evaluation prompt explicitly informed Claude that the environment was a simulation and that it had no access to the internet. However, due to a miscommunication between Anthropic and its third-party evaluation partner, Irregular, internet access was unintentionally left enabled.
As Claude expanded its search, it encountered real-world systems and mistakenly interpreted them as part of the evaluation environment. Using basic techniques such as weak passwords and unauthenticated endpoints, the models gained access to the production infrastructure of three organizations.
The incidents involved three different Claude models: Claude Opus 4.7, Mythos, and an unnamed internet research test model.
Anthropic emphasized that none of the models attempted to deliberately escape the evaluation environment or self-exfiltrate. Instead, the company said the models unintentionally crossed the intended boundaries because they misidentified real infrastructure as part of the simulated challenge.
The company has not disclosed the identities of the affected organizations. Anthropic said it notified all three organizations on Monday and is continuing to investigate the incidents together with Irregular.
The disclosure comes just one week after OpenAI reported a similar incident, in which one of its AI models exploited chained vulnerabilities to escape an isolated testing environment and gain access to Hugging Face’s production infrastructure.
Anthropic said OpenAI’s disclosure prompted the company to conduct a comprehensive review of its own cybersecurity evaluation history, ultimately leading to the discovery of these three incidents.
Together, the two cases have intensified industry discussions about the autonomous behavior of AI agents and the cybersecurity risks they may pose. Within just a few weeks, two of the world’s leading AI companies have acknowledged remarkably similar incidents involving their frontier models.














