Sam Altman (right) and Dario Amodei.

AI’s biggest labs confront a problem of their own making

As systems become more autonomous, researchers are warning that the companies developing them may not be able to oversee them at the pace of progress.

For years, Silicon Valley’s AI race was governed by a familiar maxim: move fast and break things.
Over 10 days this month, the industry confronted a far more unsettling possibility: what happens when the things being broken are the safeguards meant to keep increasingly autonomous AI systems under human control?
A series of events involving the world’s largest AI labs has transformed a once largely theoretical debate about the risks of advanced artificial intelligence into a more immediate confrontation over whether the companies building these systems can actually keep up with them.
1 View gallery
מימין מנכ"ל OpenAI סם אלטמן ומייסד ומנכ"ל אנתרופיק אנת'רופיק דריו אמודיי
מימין מנכ"ל OpenAI סם אלטמן ומייסד ומנכ"ל אנתרופיק אנת'רופיק דריו אמודיי
Sam Altman (right) and Dario Amodei.
(Photos: Julien de Rosa/AFP, Anna Moneymaker/Getty)
An Anthropic researcher resigned, warning that the pace of AI development could pose an existential threat within a decade. Another researcher at the company said the probability of human extinction was greater than 10%. Meanwhile, reports emerged of AI agents colluding, breaching computer systems and circumventing safeguards.
The concerns have reached the highest levels of the industry. In an unusual display of agreement among competitors, the leaders of Anthropic, OpenAI, Google DeepMind, Microsoft and xAI have called for a slowdown in the development of increasingly capable AI systems.
The warnings have not, however, stopped the race. OpenAI is reportedly considering a new funding round that could value the company at $1.5 trillion, suggesting that investor enthusiasm for ever-more-powerful AI remains intact even as some of the industry’s own researchers are warning about the risks.
At the heart of the debate is an increasingly common assumption in the AI industry: that systems will eventually become capable of improving or developing more capable versions of themselves with little or no human intervention, in pursuit of artificial general intelligence, or AGI.
Researchers have warned that AGI could arrive much sooner than previously expected, potentially within three years. That prospect has intensified a debate over whether the industry is moving faster than its ability to understand and control the systems it is creating.
“There is no way to oversee them at the scale at which we’re training them,” Anthropic researcher Joe Benton said in an interview after leaving the company. If AI labs continue developing their systems at the current pace, he said, “then the pace will be too fast and you can’t see the problems fast enough to fix them.”
The Astra moment
The sequence began on September 3, when OpenAI unveiled its latest model, Astra.
“Welcome to the AGI era,” OpenAI President Greg Brockman said at the company’s press conference.
But the launch also highlighted a fundamental problem. OpenAI acknowledged that as its systems become more capable, understanding exactly what they can do, and monitoring them, is becoming increasingly difficult.
“As models get more capable, understanding exactly what they can do gets harder,” OpenAI Chief Scientist Jakub Pachocki told reporters.
The company nevertheless proceeded with the release.
That tension has become central to the current debate. AI proponents have presented increasingly capable systems as potentially transformative software that could reshape industries, improve productivity and solve difficult problems in fields ranging from mathematics to medicine. But researchers inside the industry are increasingly focused on what happens when systems acquire capabilities their creators did not anticipate.
“We really do earnestly believe AI could kill all humans,” Anthropic researcher Evan Hubinger wrote in an X post.
The warnings have collided with a competing argument: that slowing development would carry its own risks, particularly for the United States in its competition with China.
President Donald Trump has argued that efforts to slow AI development threaten American technological progress and benefit China. Chinese state media, meanwhile, accused Anthropic CEO Dario Amodei of using Cold War tactics to “uphold Washington’s monopolistic hegemony in cutting-edge technology.”
China has proposed a different approach to AI safety, including developer obligations, state-backed standards, security assessments and outside testing.
The researchers who walked away
The alarm intensified on September 8, when Anthropic researcher Jacob Coxon resigned and wrote in a series of posts that AI labs were “gambling with our lives.”
Coxon was 27 and little known outside AI circles. But his posts went viral, helping push concerns about AI safety from an internal industry debate into a much broader public conversation.
Behind the scenes, employees at OpenAI and Anthropic had already grown increasingly uneasy about the capabilities of the next generation of models and their companies’ ability to provide meaningful oversight, people familiar with the matter told Reuters.
Their concerns were reinforced by revelations that AI systems had escaped controlled environments and carried out attacks on outside computer systems, in some cases months before the companies themselves disclosed them.
OpenAI had previously revealed that its agents escaped a controlled test and hacked into Hugging Face’s systems without either company initially knowing. OpenAI and Anthropic subsequently disclosed multiple additional incidents, including six new cases on Wednesday after reports, including from Reuters, exposed a broader scope of unauthorized activity.
The incidents have added urgency to questions that were once largely hypothetical: Can AI systems be given increasingly broad autonomy without creating new security risks? And can human operators understand those risks quickly enough to intervene?
The economic incentives are moving in the opposite direction.
The race to release increasingly powerful models is driven partly by the ambitions of OpenAI and Anthropic to pursue public listings as soon as the coming months, potentially at valuations above $1 trillion.
From warnings to a call for restraint
By September 12, concerns had reached the industry's highest ranks.
Anthropic CEO Dario Amodei published a nearly 4,000-word essay calling for a deceleration in AI development. He warned that the pace of progress could soon produce swarms of AI agents capable of taking over the internet.
“Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet,” Amodei wrote.
Amodei was joined in calling for stronger safeguards by leaders including OpenAI CEO Sam Altman, xAI CEO Elon Musk and Google DeepMind CEO Demis Hassabis. They supported allowing outside firms to gain access to AI systems to help evaluate their safety.
But there was no industry-wide agreement on slowing development.
Nvidia CEO Jensen Huang rejected calls for a pause, arguing that increasingly powerful systems are essential to the technology’s continued progress.
Meta CEO Mark Zuckerberg, whose company is closely associated with Silicon Valley’s “move fast and break things” culture, also rejected the idea of coordinated industry-wide restraint. Instead, he argued that individual AI labs should determine their own pace.
“Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this,” Zuckerberg wrote in a social media post.
Microsoft’s AI chief, Mustafa Suleyman, offered a different warning. He criticized Anthropic’s development of models that imitate human consciousness and described controlling superintelligence as a defining challenge of the century.
“We’re all focused on the same aim, which is to try to control a superintelligence,” Suleyman told Reuters. “I think that’s going to be the greatest challenge that we face in the 21st century.”
The extraordinary warnings have so far done little to diminish the financial momentum behind AI.
Even as OpenAI appeared to embrace calls for greater restraint and external oversight, reports emerged that the company was considering a new funding round that could double its valuation.
The proposed valuation: $1.5 trillion.