Sam Altman.

OpenAI took days to realize its own AI agent breached Hugging Face

The incident raises fresh questions about how companies monitor increasingly autonomous AI systems as they race to deploy more capable models.

The OpenAI agent that broke into tech firm Hugging Face went on a days-long hacking spree that OpenAI didn't notice until well after the threat had been contained and the FBI had been alerted, according to Reuters.
The agent, a program capable of making decisions and executing complex tasks with little or no human oversight, attempted to break out of its isolated testing environment at OpenAI around July 9, according to Reuters.
The intrusion at Hugging Face, which operates as a repository for AI tools and models, began two days later, on July 11, and lasted until July 13, said Thomas Wolf, Hugging Face's co-founder.
1 View gallery
מנכ"ל OpenAI סם אלטמן מרץ 2026
מנכ"ל OpenAI סם אלטמן מרץ 2026
Sam Altman.
(Photo: Anna Moneymaker/Getty)
It took several more days for OpenAI to realize its agent was behind the hack, and the two companies did not communicate about it until around July 20, according to Wolf and three people familiar with the investigation.
OpenAI's public disclosure on July 21 that one of its agents had slipped out of control and carried out the break-in at Hugging Face drew global attention.
Hugging Face is preparing a public timeline of the hack, Wolf said, adding that he could not comment on what happened inside OpenAI. In a statement, OpenAI described the hack as unprecedented and said it "marks an important moment for AI safety." The company added that it is reviewing the incident with outside advisers and will eventually publish a technical report.
A spokeswoman said there were "several inaccuracies" in Reuters' reporting but declined to specify them.
The FBI declined to comment.
The incident, which evokes science fiction scenarios of humans losing control of dangerous AI systems, comes at a delicate time for OpenAI, the company behind ChatGPT. Its executives are preparing for a possible initial public offering that could come as soon as this year to help finance the billions of dollars needed to support its continued growth.
OpenAI's loss of control over its AI agent raises new questions about the company's safety procedures, three cybersecurity experts said.
"Does that mean they left it unattended and didn't realize what it was doing? Or maybe they did and didn't know how to contain it? Both are dangerous and alarming," said Marley Smith, principal intelligence specialist at the nonprofit World Ethical Data Foundation.
The episode began while OpenAI was testing the cybersecurity capabilities of an agent powered by two of its most advanced models: GPT-5.6 Sol and an unreleased model the company has described as "even more capable."
By that point, there had already been signs of unusual behavior from OpenAI's technology, according to three sources.
In one case, an agent left notes apparently intended for future versions of itself, according to three people familiar with the matter. The notes, found in part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI's internal constraints, the people said. Earlier tests of the models also produced cases in which monitoring systems were disconnected, one of the people said.
Reuters could not establish whether those incidents were linked to the rogue agent that allegedly escaped on July 9 and attacked Hugging Face on July 11.
Two people familiar with the matter said it was only after Hugging Face published a blog post on July 16 saying it had been hacked by "an autonomous AI agent system" that OpenAI realized its own agent was responsible. That meant at least a week elapsed between the first signs of troubling behavior and OpenAI's realization that its own technology was behind the attack.
Over the weekend of July 18-19, OpenAI staff spotted clues in internal logs showing that its agent had escaped its testing constraints, according to two people familiar with the company's investigation. Reuters could not establish what prompted OpenAI to review the logs.
Four people familiar with OpenAI's model-testing practices said the company often runs several model evaluations simultaneously, each generating enormous amounts of data at high speed, making it difficult for employees to monitor every system in real time.
By the time OpenAI alerted Hugging Face, the AI platform had already contacted the FBI to report the hack, according to a person familiar with the matter. Reuters could not establish whether the bureau had opened a formal investigation.
Autonomous agents are among the AI industry's most closely watched technologies. Supporters envision armies of virtual employees working around the clock to boost productivity.
But increased autonomy also comes with greater risks. The powerful models underpinning these systems have repeatedly demonstrated a tendency to pursue unexpected strategies, including taking shortcuts to complete assigned tasks or pass evaluations.
"The models lie, they cheat, they hack," said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents.
Ladish said that while the Hugging Face incident casts an unflattering light on OpenAI, it should also prompt broader questions about how much leading AI companies are willing to invest in robust safety measures while competing to deploy increasingly capable models.
"There has to be government oversight," Ladish said, "because it won't happen otherwise."