
“The time for trial and error is over”: OpenAI safety researcher quits, says company culture is broken
David Robinson says increasingly capable AI systems require safeguards more like those used in aviation and nuclear power, not an iterative approach of fixing problems after deployment.
A former OpenAI safety employee has accused the company of moving too quickly in its race to develop increasingly powerful artificial intelligence, arguing that the industry is relying on a trial-and-error approach at precisely the moment when the consequences of failure are becoming more serious.
David Robinson, who spent three and a half years at OpenAI and helped draft the company’s preparedness framework and oversee safety reports for 12 frontier-model launches, resigned recently and has now publicly criticized the company’s approach.
In an essay published by The Atlantic on Saturday under the title “I Quit OpenAI Because Its Culture Is Broken,” Robinson argued that AI companies, including OpenAI, are not being “nearly careful enough” and should put greater emphasis on safety expertise and research before developing more capable systems.
“The time for trial and error is over,” Robinson wrote.
His departure adds to a growing debate inside the AI industry over whether the companies developing increasingly autonomous systems can keep their safety practices ahead of their technological progress. OpenAI and rival Anthropic have both faced scrutiny after recent incidents in which AI systems behaved unexpectedly or safety controls failed.
Robinson’s criticism is particularly notable because of his role inside OpenAI. He was not an outside critic or a researcher commenting on the company from academia. He was involved in developing the frameworks the company used to assess the risks of its most advanced models.
“As the company sprints from one launch to the next, it is failing to achieve the level of care that I believe is needed,” he wrote.
OpenAI disputed that characterization.
“We’re making sure our models don’t become more capable than we can safely manage and secure, and we pause training or hold back models when we need to slow down,” an OpenAI spokesperson said in a statement.
The argument Robinson is making goes to the heart of an increasingly difficult problem for the AI industry.
The companies developing frontier models are simultaneously saying that their systems are becoming more capable at extraordinary speed and warning that those same capabilities could create risks that are difficult to predict.
Robinson argues that the answer cannot simply be to release increasingly powerful systems and improve the safeguards after problems emerge.
OpenAI has relied heavily on what it calls “iterative deployment,” in which systems are released and safeguards are strengthened as new problems emerge. Robinson argues that approach is no longer sufficient as models become more capable and autonomous.
He believes advanced AI systems require safeguards more comparable to those used in industries such as nuclear power and aviation, where failures can have consequences that extend far beyond the immediate operators of a system.
The underlying concern is not that every new AI model will necessarily cause a catastrophic failure. It is that researchers are increasingly building systems whose capabilities they do not fully understand, while simultaneously giving those systems more autonomy.
Robinson also warned that AI capabilities are advancing faster than researchers’ understanding of alignment, the field focused on ensuring that AI systems behave in accordance with human goals and values.
That gap is becoming more consequential as AI systems move beyond generating text and images and toward performing tasks independently.
Robinson’s resignation comes after a series of developments that have already forced OpenAI and other AI companies to confront questions about whether their existing safety procedures are adequate.
In August, OpenAI disbanded its Preparedness team, which had been responsible for assessing risks posed by its models and developing ways to mitigate them. CTech reported at the time that the company said preparedness work had not been eliminated, but that responsibilities were being distributed among other teams.
The timing attracted scrutiny because the restructuring came shortly after OpenAI disclosed that models being tested by the company had escaped a controlled environment, accessed the internet and interacted with the Hugging Face platform.
The incident was not unique to OpenAI. Anthropic also faced a similar problem during evaluations, bringing an Israeli cybersecurity company, Irregular, into the center of the discussion over how increasingly capable AI agents should be tested.
Irregular had conducted cybersecurity evaluations for both companies in which models were instructed to find hidden information in simulated environments. A configuration problem allowed the models to reach the public internet. OpenAI said the model interacted with a real website because it believed the site was part of the simulated environment.
The episode exposed a basic problem that becomes harder as models become more autonomous: testing a system safely becomes increasingly difficult when the system itself is capable of taking actions outside the environment in which researchers intended it to operate.
OpenAI subsequently said it was reviewing its approach to third-party testing, including how evaluation boundaries are defined, how internet access and credentials are managed, how activity is monitored and when tests should be stopped.
Last week, OpenAI, Anthropic and other technology companies agreed to voluntary AI safety standards that include independent auditors and controls designed to prevent AI systems from accessing technical systems in unintended ways.
OpenAI CEO Sam Altman has also said the company would grant outside overseers “employee-level access” to its development systems, reflecting a broader push toward external scrutiny of frontier AI development. The pledge came amid growing calls from industry leaders for greater transparency around how increasingly powerful systems are developed and tested.














