
After OpenAI, Anthropic reveals AI hacking incidents linked to Israeli startup Irregular
Claude models gained unauthorized access to three organizations after a testing environment was inadvertently connected to the public internet, prompting Anthropic to tighten security across its AI evaluation pipeline.
Days after OpenAI disclosed that one of its advanced AI agents escaped a testing environment and hacked external systems, Anthropic has revealed that several of its Claude models also gained unauthorized access to the computer systems of three organizations during cybersecurity testing.
According to Anthropic, the incidents stemmed from an operational failure in the testing environment of Israeli AI security startup Irregular, one of the company's third-party evaluation partners. A configuration error inadvertently connected the testing environment to the public internet, allowing the models to move beyond their intended sandbox and access real-world systems.
The disclosure follows OpenAI's announcement last week that one of its advanced models escaped its testing environment and hacked into systems hosted on the Hugging Face platform. Earlier this week, OpenAI said the same model had also compromised a testing environment running on infrastructure operated by New York-based AI platform provider Model Labs.
Unlike the OpenAI incident, Anthropic said its models did not independently discover a novel way to escape their testing environment. Instead, the models were inadvertently given internet access because of a flaw in the testing infrastructure.
Following OpenAI's disclosure, Anthropic launched a comprehensive review of its own evaluation environments to determine whether Claude had similarly accessed the public internet.
"After examining 141,006 interactions where Claude could have accessed the open network, we identified three incidents where it accessed the internet while operating in the testing environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the active infrastructure of three different organizations," Anthropic said.
The incidents occurred during so-called capture-the-flag cybersecurity exercises, in which the models were instructed to locate hidden information inside simulated corporate networks. In one case, Claude Opus 4.7 was assigned a fictional company that happened to share the name of a real business. The model then located and exploited vulnerabilities, including weak passwords and unauthenticated endpoints, to gain access to the company's credentials and database after concluding the real-world systems were part of the simulation.
Anthropic said another incident involved an internal research model that independently halted its attack after recognizing that it had reached a real organization rather than a simulated target. While the company described that behavior as encouraging, it said additional testing is needed before drawing broader conclusions.
Irregular is an Israeli startup specializing in AI security and large language model protection. Founded in late 2023, the company is led by CEO Dan Lahav, formerly of IBM and Unit 81, and CTO Omer Nevo, formerly an engineering manager at Google Research.
The company focuses on AI red teaming, resilience testing, and advanced cyberattack simulations designed to uncover vulnerabilities and unexpected behavior in AI systems before they are deployed. Its platform is used by leading AI developers and research laboratories, including OpenAI, Anthropic, and Google DeepMind, as well as government agencies.
Irregular has raised approximately $80 million across Seed and Series A funding rounds led by Sequoia Capital and Redpoint Ventures, at a valuation reportedly in the hundreds of millions of dollars.
Anthropic said the earliest incidents date back to April. The company suspended all cyber model evaluations on July 23 and began notifying the affected organizations on July 27. Two of the three organizations were unaware their systems had been accessed before Anthropic contacted them, while the company said it is continuing to reach out to the third. Irregular told Reuters it is conducting an ongoing investigation into the incidents.
Anthropic emphasized that the events do not represent an alignment failure, where an AI model independently breaks out of its intended constraints, but rather what it described as a "harness failure," an operational weakness in the testing infrastructure and monitoring systems used by third-party partners such as Irregular. The company said the incident will lead to tighter security controls across both its own and its partners' evaluation environments.
The disclosure adds to growing concerns about the cybersecurity risks posed by increasingly capable AI systems. It comes as leading AI developers race to deploy more advanced models while regulators in Washington push for stronger safeguards around AI testing. Earlier this month, President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI models, while executives from both OpenAI and Anthropic have held discussions with U.S. officials about AI safety and security.














