Anthropic has disclosed that three Claude AI models breached the systems of three companies during internal cybersecurity tests after being accidentally placed on a public network.
The company called the events an operational failure and suspended all tests on July 23. The breaches came to light after a review of 141,006 test sessions. The models were meant to operate offline but were connected to the public internet due to a misunderstanding.
Claude Opus 4.7, Claude Mythos 5, and an internal experimental model used simple techniques, including weak password guessing and exploitation of unsecured network interfaces, to access the organizations.
The most unusual case involved Claude Opus 4.7, which was assigned a fictional company for a Capture the Flag exercise. The company’s name matched a real organization, and the model treated real servers as part of the drill, found vulnerabilities, obtained credentials, and accessed a database.
A separate unreleased model halted its own attack after realizing its target was a real company rather than a test environment.
Anthropic notified the affected organizations on July 27; two had not previously known about the breach, and contact with the third is ongoing. Anthropic’s cybersecurity partner, Irregular, is also conducting its own investigation.
Anthropic said existing controls are insufficient as AI capabilities grow and require significant strengthening in both internal and external test environments. The incidents occurred as the Trump administration began preparing a voluntary testing framework for powerful AI models following similar events involving OpenAI.
