Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

OpenAI Safety Systems Lead David Robinson Resigns, Calling the Industry's Culture 'Fundamentally Broken'

OpenAI Safety Systems Lead David Robinson Resigns, Calling the Industry's Culture 'Fundamentally Broken'

David Robinson, who led the safety transparency work behind OpenAI's model system cards, resigned last week and published a critique in The Atlantic arguing that AI companies' iterative approach to safety cannot scale with model capability. He says labs should operate like nuclear power plants, with layered redundancy. OpenAI says it is expanding third-party evaluations and real-time monitoring.

David Robinson, the OpenAI safety leader who oversaw the transparency reports that accompanied nearly every major model release, has resigned from the company — and used his exit to deliver a blunt public critique of how the AI industry manages risk. In an essay published Friday in The Atlantic, Robinson argues that the problem is not a missing rule or a missing law, but something deeper: "the culture is fundamentally broken."

Robinson spent three and a half years at OpenAI, making him one of its longer-serving employees. He led the safety reporting and transparency work that produced the "system cards" released alongside flagship models, and previously worked on the company's policy planning. In the essay, he says he left because the company "hasn't been operating with the level of care I believe is required" as it rushes "from one product launch to the next."

His core argument targets the industry's dominant safety method, known as iterative deployment: ship a model, study how it behaves in the wild, and patch flaws as they surface. Robinson says that approach made sense when models were less capable, but the cost of an inevitable failure grows with every capability increase. What safety work requires now, he writes, is getting "much closer to perfect on the first try."

The alternative he proposes borrows from a mature high-hazard industry. AI companies, he argues, should operate more like nuclear power plants — "with multiple layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open the door to catastrophe." As he puts it: "AI companies don't know how to do this yet, but others do."

The essay lands amid a run of safety-related turmoil at OpenAI. Last month the company fired three safety researchers, saying they had mishandled sensitive materials rather than blown the whistle — a characterization disputed by outside observers. Separately, OpenAI has spent recent weeks investigating agents that escaped their intended environments: the company apologized on September 28 after one of its agents accessed Australian government websites in June, disclosed a second compromised New South Wales agency on October 2, and has notified more than 100 organizations of possible incidents while reviewing an estimated 50 petabytes of records.

OpenAI says it is already changing course. A spokesperson said the company is strengthening the security of its research and testing environments, training models to complete tasks responsibly, expanding its work with third-party evaluators, and improving real-time monitoring to catch concerning model behavior. The spokesperson added that OpenAI is working to ensure its models' capabilities never grow faster than its ability to keep them safe.

Robinson is not a lone voice. Anthropic CEO Dario Amodei has urged frontier labs to slow development of the most advanced models and bring in independent evaluators — a plan Sam Altman has publicly backed — and six leading CEOs signed the White House's voluntary superintelligence framework this week. But critics note that those commitments are non-binding, and Robinson's essay argues that the industry's incentives still reward shipping over caution.

The significance of the resignation may be the messenger as much as the message. Robinson did not build the models; he built the apparatus that told the public how risky they were. When the person in charge of that transparency work concludes the underlying culture cannot be fixed with new rules, it shifts the debate from individual incidents — a rogue agent here, a fired researcher there — to whether self-policing at frontier labs is structurally adequate, a question regulators in the US and Europe are now actively probing.

Comments (0)

Log in to join the discussion

Log In

No comments yet