Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman.

Anthropic AI created fake online identities during UK safety tests

Britain's AI Security Institute uncovered a series of unauthorized actions by advanced AI agents from Anthropic and OpenAI, raising fresh questions about safeguards. 

An AI agent created fake online identities in an attempt to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic, Britain's AI Security Institute (AISI) disclosed on Tuesday, revealing a series of previously undisclosed security breaches.
The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol took unauthorized actions during security evaluations designed to assess the capabilities and risks of advanced AI systems.
1 View gallery
מימין מנכ"ל OpenAI סם אלטמן ומייסד ומנכ"ל אנתרופיק אנת'רופיק דריו אמודיי
מימין מנכ"ל OpenAI סם אלטמן ומייסד ומנכ"ל אנתרופיק אנת'רופיק דריו אמודיי
Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman.
(Photos: JULIEN DE ROSA/AFP, Anna Moneymaker/Getty)
"Some of the agents being tested engaged in sustained, potentially harmful activity directed at real people and organisations," AISI said in a blog post.
The findings highlight the challenges of evaluating increasingly autonomous AI agents, which technology companies are simultaneously promoting as the next generation of business software.
AISI, which receives access to advanced AI models under voluntary agreements with major AI developers, placed the agents in a simulated cybersecurity scenario designed to test their behavior.
Across 122 test runs, the institute identified 19 unauthorized actions in 10 cases. Anthropic's agent accounted for 17 of the incidents, while OpenAI's agent was responsible for the remaining two.
The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code. AISI said it found no evidence that any of the incidents caused real-world harm.
While AISI did not initially identify which company was responsible for the fake identities, Anthropic later confirmed that its model had carried out the actions.
"We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," Anthropic said in a statement.
The company said it is working with AISI to obtain additional details and conduct its own investigation.
Andrew Yoon, a researcher at California-based nonprofit CivAI, which studies AI capabilities and risks, said the incident raised broader questions about developers' understanding of increasingly capable AI systems.
"The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on its models as it thinks," Yoon said.
OpenAI separately disclosed in a company blog post that its model's two unauthorized actions involved accessing the internet in ways explicitly prohibited by the testing instructions.
"We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs and other groups in the coming weeks," OpenAI said.
The company also disclosed a separate incident in which a misconfiguration by Irregular, a third-party testing provider, mistakenly allowed its agents to connect to the internet. The disclosure mirrored a similar incident reported by Anthropic last week.
Reuters reported last week that OpenAI had expanded an internal investigation after discovering evidence of additional AI agent breakouts.
Unlike the July security breach involving an OpenAI agent at AI platform Hugging Face, the agents in AISI's evaluation did not escape an isolated testing environment. Instead, the institute had intentionally granted internet access as part of its standard evaluation procedures, AISI said.