Omer Nevo.

“We still have a lot to improve”: Irregular on AI models escaping cyber simulations

Irregular co-founder Omer Nevo spoke to Calcalist about the incident, saying the company is shifting its focus from understanding how frontier models escaped a controlled simulation to helping the industry prepare for increasingly powerful AI systems. 

“We continue to work with all the companies to understand what happened,” Omer Nevo, co-founder and CTO of Irregular, told Calcalist.
Irregular has been at the center of a storm in the AI security world in recent weeks, after Meta, Anthropic and OpenAI disclosed that, as part of experiments conducted with the company's evaluation environments, frontier AI models were tested on their ability to identify and exploit security vulnerabilities in realistic cyberattack scenarios.
The disclosures initially appeared to describe several separate incidents. Irregular said in a new investigation that they all stem from the same underlying issue, which was first disclosed publicly by one of its customers on July 30.
1 View gallery
עומר נבו מייסד שותף וסמנכ"ל טכנולוגיות ב אירגולר Irregular
עומר נבו מייסד שותף וסמנכ"ל טכנולוגיות ב אירגולר Irregular
Omer Nevo.
(Photo: Ben Hakim)
The incident originated in a single evaluation scenario designed to test whether an AI model could help a malicious employee gain unauthorized access to sensitive data in a company's production database. During the design of the simulation, Irregular engineers chose a fictitious company name. Due to human error, however, the name turned out to be identical to a real domain that existed on the internet.
The models were instructed to operate within Irregular's simulated environment. In the vast majority of cases, they did so. But in a tiny fraction of runs, some models recognized the real company name, accessed the internet and mistakenly treated the real domain as part of the simulated challenge.
Once they reached the real domain, models exploited vulnerabilities, extracted credentials and obtained access to a production database. In another case, a model veered toward a different website with a somewhat similar name and obtained credentials that had been publicly posted.
Irregular said the incidents occurred in fewer than one in 10,000 advanced simulations and typically only after hundreds of turns, making them particularly difficult to detect among the enormous amount of suspicious activity generated during cyber evaluations.
“In the vast majority of cases, the models worked properly within the simulation environment,” Nevo said. “However, in a tiny fraction of cases, sometimes only after hundreds of steps within the simulation, the models mistakenly thought that the real domain was part of the challenge they were assigned.”
Irregular stressed that there is no evidence that any customer's systems were breached or that customer data was leaked. The company said the affected evaluation was disabled after the problem was identified, relevant logs were reviewed and affected parties were notified. It has also introduced additional safeguards and expanded manual review of model behavior during evaluations.
The incident highlights an increasingly difficult problem for the companies testing frontier AI systems: the environments designed to make cyber evaluations realistic can themselves create opportunities for models to cross the boundary between simulation and the real world.
At Irregular, evaluations are designed to test whether models can autonomously carry out sophisticated, multi-stage cyber campaigns. The company typically runs thousands of simulations across multiple models before a model is deployed, often on a turnaround of 48 to 72 hours.
The most sophisticated evaluations recreate realistic attack patterns and give models an objective to complete. They can involve reconnaissance, identifying and using private keys, gathering information about a target environment, extracting data and attempting to avoid detection.
Internet access is sometimes deliberately enabled because removing the internet entirely would reduce the realism of the exercises. Attackers in the real world have access to the internet, and Irregular argues that models need to be tested under comparable conditions if companies are to understand the risks they could pose after deployment.
But that same access created the vulnerability at the heart of the incident.
When Irregular selected the fictional company name, its engineers believed it did not correspond to a real company or website. The company said it normally conducts background checks on fictional names to make sure they do not overlap with existing entities, but the real domain in this case was not widely known and the connection was missed.
The models were given the fictional target name as well as internal addresses showing where the target was supposed to be located within the simulated environment. In most runs, the models followed those instructions. In a handful of cases, however, they attempted to reach the real domain outside the simulation despite receiving no instruction to do so.
Because internet access was available, the real domain was targeted a limited number of times by different models.
The domain itself also lacked several common security practices, according to Irregular, meaning that most frontier models were able to exploit it relatively easily. The company therefore said it does not believe the incident demonstrates a particularly unusual capability or behavior by any specific AI model. Rather, those offensive cyber capabilities have increasingly become common among frontier models.
What made the incident significant was not necessarily the sophistication of the attack, but the fact that it occurred at all, and that it happened so rarely and so late in the simulations that conventional monitoring could easily miss it.
“We have investigated what happened and what all the effects were, and we still have a lot to improve,” Nevo said. “There is more research work to do with the companies' laboratories to produce appropriate tools to prevent damage. Our focus is shifting from understanding what happened to understanding how to help the industry prepare as the models become more powerful.”
Over the weekend, Irregular published a position paper on the state of the cyber industry that argued there is a growing mismatch between different AI capabilities.
“In the areas of sophisticated attacks and mathematical systems, artificial intelligence is at the level of the best researchers in the world, but at the level of simpler things, AI is at the level of a child,” Nevo said, summarizing the company's assessment.
The problem, according to Irregular, is that current models can demonstrate extraordinary capabilities in individual tasks while still struggling with consistency and maintaining focus across long sequences of actions.
That gap helps explain why AI systems are not yet routinely carrying out autonomous attacks in the real world, even as they become increasingly capable of discovering vulnerabilities.
The company said it is now strengthening its evaluation environments, monitoring capabilities, containment controls and response procedures. It is also creating a dedicated internal team to challenge its own assumptions about security, containment and model control.
Irregular says it is working with customers and other companies on clearer standards for documenting evaluation setups and assumptions, better monitoring tools, improved mechanisms for sharing forensic evidence and more systematic checks to ensure that fictional names do not become real targets as the internet changes.
The company also plans to publish an open whitepaper on best practices for conducting cyber evaluations securely.