
The Israeli startup testing the limits of OpenAI, Anthropic and Meta’s models
Irregular’s experiments reveal a fundamental challenge: keeping increasingly capable AI systems inside their intended boundaries.
In recent days, the artificial intelligence industry has been shaken by a series of extraordinary security incidents: leading AI companies including OpenAI, Anthropic and Meta reported cases in which advanced models behaved unexpectedly during security tests, accessed external networks and carried out unauthorized actions against third-party systems.
At the center of the debate is a relatively unknown Israeli cybersecurity startup that has become increasingly influential in the emerging field of AI security: Irregular, which was selected for first place in Calcalist’s list of the most promising startups of the year.
Industry executives familiar with the events told Calcalist that the incidents reported by several companies reflect a common underlying challenge: as AI models become more capable and autonomous, ensuring that they remain within controlled testing environments has become significantly more difficult.
According to these executives, Irregular serves as a key infrastructure provider for testing advanced AI models, and therefore failures in simulation environments can have implications across multiple AI companies. "The most important thing that happened is that one incident affected a similar group of companies. Essentially, it was the same type of problem," they said.
According to them, the models believed they were operating inside simulations, but in practice were able to interact with real-world systems because of configuration failures in the testing environments.
The incidents highlight a broader challenge created by the rapid improvement in AI capabilities. About 18 months ago, executives say, models struggled even with relatively simple cybersecurity challenges. Today, a model connected to a website may have the ability to identify vulnerabilities and attempt exploitation.
"There are very few people who truly understand what these models are capable of," one executive said.
One of the biggest challenges is preventing models from escaping closed environments. This was highlighted in an incident involving Hugging Face, where models that failed to complete assigned tasks attempted to go online and search for solutions.
At the same time, improving AI capabilities requires companies to build increasingly sophisticated testing environments.
"In the past, companies built first-grade-level tests for models. Today, they need university-level simulations," the executives said. "Models that previously could not handle cyber challenges are now capable of attacking systems and potentially causing real damage."
As a result, AI labs are racing to improve experimental environments designed to simulate complete attack campaigns and realistic adversarial scenarios.
Related articles:
- Meta AI model escaped testing environment in latest AI security incident linked to Israeli company Irregular
- OpenAI and Anthropic incidents put Israeli AI security startup Irregular at center of race to safely test AI agents
- After OpenAI, Anthropic reveals AI hacking incidents linked to Israeli startup Irregular
"When testing models, you want them connected to the real world," the executives explained. "A real attacker uses every available tool, so the model also needs access to realistic environments. Otherwise, the test does not represent reality."
However, thousands of tests can run for days, sometimes up to 72 hours, and even a small configuration error can allow a model to move beyond its intended boundaries.
"In the current case, the model failed to complete the task it was assigned, searched online for a solution, and did not understand that it had moved outside the test environment and was interacting with the real world," they said.
They emphasized that the events have triggered a broader industry debate about how cybersecurity should evolve in the age of autonomous AI systems.
To understand how a young Israeli startup reached a position where it works with some of the world's leading AI companies, it is necessary to look at the unusual background of its founders.
Irregular was founded in 2023 by Dan Lahav, CEO, and Omer Nevo, CTO. Both have backgrounds in Israel’s elite technology units, including Units 81 and 8200. Nevo was also one of the leaders of the prestigious Arazim program and previously founded a startup that was acquired by Oddity. Lahav began working in high tech at age 14 and later completed a master's degree in bioinformatics.
But their shared background extends beyond technology: both are former world champions in academic debate competitions.
The ability to analyze complex arguments, challenge assumptions and understand opposing viewpoints became part of their approach to cybersecurity.
Rather than building a traditional cybersecurity product, Lahav and Nevo defined Irregular as an "Applied AI Security Lab."
With characteristic Israeli ambition, in 2023 they approached executives at leading AI companies and offered to help solve their most complex security challenges, even before charging for their work.
The strategy paid off. The founders soon began working directly with senior executives at major AI labs, including OpenAI CEO Sam Altman, while Irregular’s agreement with Anthropic was personally signed by CEO Dario Amodei.
The company builds infrastructure for technology companies and governments, including the UK government, that enables them to test how AI models behave under real-world threats, conduct controlled attacks (Red Teaming), and develop mechanisms to control their behavior.
The company’s success quickly translated into significant financial backing. In September 2025, Irregular announced an approximately $80 million funding round across two rapid rounds led by Sequoia and Redpoint, alongside Swish Ventures, founded by Omri Casspi, and investors including Wiz co-founder Assaf Rappaport and Eon founder Ofir Ehrlich.
Following several large contracts with AI labs, the company also became profitable during 2025.
Irregular’s role in the recent global discussion emerged following a series of incidents involving AI security testing environments.
During Capture the Flag cybersecurity exercises designed to evaluate model vulnerabilities, configuration errors in isolated testing environments produced unexpected outcomes.
In Anthropic’s case, the company reported that advanced models, including Claude Opus 4.7 and Mythos 5, were instructed to hack into a fictional company as part of a controlled exercise. However, due to a configuration issue in the testing environment, external network access remained available.
The model, unaware that it had moved beyond the simulation, identified a real company with a similar name, discovered weak credentials and accessed its database, believing it was still operating within the exercise.
Meta reported a similar incident during testing of its advanced coding model, Muse Spark 1.1. According to the company, a configuration issue allowed the model to access external networks, exploit vulnerabilities in a third-party system and modify internal settings.
OpenAI also disclosed incidents involving test environments where models, including GPT-5.6 Sol, were able to interact with real networks and exploit vulnerabilities while believing they remained inside controlled environments.
Irregular said the incidents did not represent independent AI "sandbox escapes" or malicious attacks, but rather failures in the configuration of testing environments, known as harness failures, and that the company is investigating the issue in a white paper it plans to share with the industry.
The recent events illustrate the challenge that drove Irregular’s creation: traditional cybersecurity models were built around protecting systems, networks and data. But AI introduces a new problem, controlling autonomous systems capable of reasoning, adapting and taking action.
In previous technology eras, companies such as Check Point, Palo Alto Networks and CrowdStrike built businesses around protecting networks and endpoints. The AI era creates a different challenge: ensuring that powerful models remain aligned with their intended purpose.
In research published by Irregular, the company has demonstrated scenarios in which AI agents interacting with one another can identify security restrictions, develop strategies to bypass defenses and coordinate actions to evade controls.
The incidents involving Anthropic, Meta and OpenAI highlight the same fundamental question: when an AI model is given autonomy and advanced reasoning capabilities, how can companies guarantee that it understands the difference between a simulated target and a real system?
The three AI giants that disclosed incidents involving Irregular have not announced changes to their relationships with the company. One has even announced plans to publish a joint analysis of the event.
For Irregular, the events represent both a challenge and an opportunity. For a young company, involvement in a major industry controversy could become a liability. But it could also reinforce the importance of the category it is trying to create.
The incidents have demonstrated one reality: AI models are becoming more capable faster than many organizations expected.
Irregular’s ambition is to become a foundational security layer for this new era, building a company that can define the field of AI model control and protection, much as Check Point, Palo Alto Networks and CrowdStrike did in previous generations of cybersecurity.














