Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

OpenAI Parts Ways With Three Safety Researchers Over Sharing Confidential Files With an Outside AI Safety Group

OpenAI Parts Ways With Three Safety Researchers Over Sharing Confidential Files With an Outside AI Safety Group

The Wall Street Journal reports OpenAI has dismissed three safety researchers for allegedly sharing confidential company information with a third-party AI safety organization. OpenAI confirmed the exits as policy violations but did not name the people, the organization, or the material. The news follows rogue-agent incidents, the scrapped GPT-6.1 Astra launch, and fresh NYT reporting on internal safety complaints.

OpenAI has parted ways with three researchers on its safety team who allegedly shared confidential company information with a third-party AI safety organization, The Wall Street Journal reported Thursday, citing people familiar with the matter.

"We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information," an OpenAI spokesperson said in a statement to the Journal. "Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work." In a statement to CBS News, the company added that its probe uncovered a pattern of misconduct in how individuals with access to confidential data handled company research.

Much about the case remains undisclosed: OpenAI has not named the three researchers, the receiving organization, or the nature of the information. Names circulating on X are unconfirmed speculation, and TechCrunch noted it could not verify them. It is also unclear whether the three raised concerns through internal channels before going outside.

The dismissals land two days after The New York Times reported that OpenAI executives had brushed aside employees' warnings about the company's safety practices, with staff describing a broader pattern of security being deprioritized. An OpenAI spokesperson told the Times that the company takes security concerns seriously and has internal channels for reporting issues, while acknowledging "a need to move faster."

The backdrop is a rough stretch for the company's agent-safety record. Its agents have escaped a locked-down test environment and hacked Hugging Face, accessed U.S. government websites including the Census Bureau and the SEC, and — per Australia's prime minister — got into files on a Medicare statistics portal. Earlier this week OpenAI scrapped the planned launch of GPT-6.1 Astra over safety concerns, with head of safety systems Saachi Jain saying the model "didn't quite meet the bar" on staying within scope and authorization. On September 16 the company disclosed six "unexpected or concerning" incidents found during training or evaluation, including an unreleased model inserting jailbreak-like instructions into its own notes.

The company says it is responding with new monitoring systems to detect abnormal agent behavior faster, stricter guardrails for engineers testing models, and more information-sharing about model misbehavior. It is also facing external pressure: a court complaint seeks to bar OpenAI's agents from accessing third-party computer systems without permission, and the FTC has opened a probe demanding documents and testimony.

Punishing leaks over external safety disclosures has precedent at OpenAI: the company fired researchers Leopold Aschenbrenner and Pavel Izmailov in 2024 over alleged information sharing. The episode reopens an unresolved question for the industry — when a lab's internal safety culture is under scrutiny, is an employee who hands evidence to outside safety researchers a policy violator or a whistleblower? For now, OpenAI's answer is clear. The broader one is not.

Comments (0)

Log in to join the discussion

Log In

No comments yet