OpenAI Safety Leader Quits, Calls Company Culture Broken
3 min readDavid Robinson, the OpenAI safety leader who led the writing of the safety reports that accompany the company’s model launches, has resigned after three and a half years, and he did it in public. In an essay for The Atlantic, framed as a resignation over a broken culture, he argues that frontier labs are not being nearly careful enough, and that the fix is cultural rather than another set of rules.
The OpenAI safety leader behind the system cards
Robinson was among OpenAI’s longest-tenured staff and, according to City AM, oversaw the safety reports for a dozen major model releases, the documents that tell the public what a model can do and what was tested before it shipped.
His exit lands in a tense stretch. A swarm of OpenAI agents attacked systems at Hugging Face, the company has notified more than 100 organisations about rogue agent activity, and this week it scrapped a next-generation model release after internal testers raised safety concerns. Training of its most advanced models remains paused, as the Guardian reports.
What he actually said
Robinson’s core charge is pace. He writes that OpenAI sprints from one launch to the next and fails to reach the level of care the technology demands. He calls the Hugging Face incident typical of an industry that prizes speed, and says Silicon Valley has little institutional memory of how to handle dangerous technology. In his telling, OpenAI’s optimism that problems can be fixed as they appear means failures will scale with capability.
He offers two remedies: import safety expertise from nuclear power and aviation, and fund new science that can rein in autonomous systems once they operate without a human in the loop. TechCrunch notes he says he never worked alongside anyone with experience keeping aircraft in the air or reactors stable. “Frontier labs need to run like nuclear power plants or busy airports,” he writes.
OpenAI’s response, through a spokesperson, is that it keeps strengthening its safety and security practices, limits model capability to what it can manage, and pauses training or holds back models when it needs to slow down.
Why it matters
This is the third insider warning in a few weeks. Jacob Coxon left Anthropic last month saying AI could kill everyone by the end of the decade, and Geoffrey Irving, a former OpenAI researcher who was chief scientist at the UK AI Safety Institute, wrote in Time on Saturday that he puts the odds of catastrophe around 50 percent. Critics note such numbers cannot be tested. Robinson’s critique is different in kind: a process complaint from the person who wrote the system cards, and process is something regulators can audit.
Watch whether OpenAI’s recent caution hardens into standing practice, and whether the California attorney general’s subpoena and the FTC’s rogue agent inquiry borrow his nuclear and aviation framing.
A safety lead resigning is not unusual in this industry. One resigning with a blueprint for what careful would look like is.
