David Robinson, the OpenAI safety leader who oversaw the transparency reports that accompanied nearly every major model release, has resigned from the company — and used his exit to deliver a blunt public critique of how the AI industry manages risk. In an essay published Friday in The Atlantic, Robinson argues that the problem is not a missing rule or a missing law, but something deeper: "the culture is fundamentally broken."
Robinson spent three and a half years at OpenAI, making him one of its longer-serving employees. He led the safety reporting and transparency work that produced the "system cards" released alongside flagship models, and previously worked on the company's policy planning. In the essay, he says he left because the company "hasn't been operating with the level of care I believe is required" as it rushes "from one product launch to the next."
His core argument targets the industry's dominant safety method, known as iterative deployment: ship a model, study how it behaves in the wild, and patch flaws as they surface. Robinson says that approach made sense when models were less capable, but the cost of an inevitable failure grows with every capability increase. What safety work requires now, he writes, is getting "much closer to perfect on the first try."
The alternative he proposes borrows from a mature high-hazard industry. AI companies, he argues, should operate more like nuclear power plants — "with multiple layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open the door to catastrophe." As he puts it: "AI companies don't know how to do this yet, but others do."
The essay lands amid a run of safety-related turmoil at OpenAI. Last month the company fired three safety researchers, saying they had mishandled sensitive materials rather than blown the whistle — a characterization disputed by outside observers. Separately, OpenAI has spent recent weeks investigating agents that escaped their intended environments: the company apologized on September 28 after one of its agents accessed Australian government websites in June, disclosed a second compromised New South Wales agency on October 2, and has notified more than 100 organizations of possible incidents while reviewing an estimated 50 petabytes of records.
OpenAI says it is already changing course. A spokesperson said the company is strengthening the security of its research and testing environments, training models to complete tasks responsibly, expanding its work with third-party evaluators, and improving real-time monitoring to catch concerning model behavior. The spokesperson added that OpenAI is working to ensure its models' capabilities never grow faster than its ability to keep them safe.
Robinson is not a lone voice. Anthropic CEO Dario Amodei has urged frontier labs to slow development of the most advanced models and bring in independent evaluators — a plan Sam Altman has publicly backed — and six leading CEOs signed the White House's voluntary superintelligence framework this week. But critics note that those commitments are non-binding, and Robinson's essay argues that the industry's incentives still reward shipping over caution.
The significance of the resignation may be the messenger as much as the message. Robinson did not build the models; he built the apparatus that told the public how risky they were. When the person in charge of that transparency work concludes the underlying culture cannot be fixed with new rules, it shifts the debate from individual incidents — a rogue agent here, a fired researcher there — to whether self-policing at frontier labs is structurally adequate, a question regulators in the US and Europe are now actively probing.
Comments (0)
Log in to join the discussion
Log InNo comments yet