ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA ChatGPT Terence Tao Amplifies a Call to Boycott OpenAI After It Dumps 722 AI-Generated Math Proofs on GitHub Coding Assistants JetBrains' Mellum2.1 Goes From 2.0 to 47.0 on SWE-bench Verified: a 12B Open Model Rebuilt by Reinforcement Learning News Huawei Hubble and Lei Jun's Shunwei Back DiffuSpace: Two Rounds Total Close to 500 Million RMB, a Record for Diffusion Language Models Business USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data AI Agents Goodfire Puts Monitors Inside the Model: 94% of Malicious Agent Sessions Caught for About $51 in Company Tests Claude Anthropic Makes Cruelty Toward Claude a Policy Violation in First Usage-Policy Rewrite in Over a Year Coding Assistants Harness Buys Augment Code's Cosmos to Complete Its Autonomous Software Factory, From Ticket to Merge-Ready PR Claude Anthropic Turns Claude Into a BI Tool: Dashboards and Motion Enter Beta as Docs, Slides and Design Go GA

USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data

USA Today's Publisher Sues OpenAI for More Than $250 Million, Citing 160,000 Entries in GPT-2's Training Data

USA Today Co. and 13 affiliates sued OpenAI in Manhattan federal court, saying its articles make up more than 160,000 WebText entries and 122 million C4 tokens. The complaint seeks over $250 million and destruction of GPT models trained on content from 19 publications - and singles out ChatGPT's article summaries as substitutes for the stories themselves.

USA Today Co., the publisher of USA Today and a chain of local papers, sued OpenAI in the Southern District of New York on Thursday, accusing the company of copying hundreds of thousands of articles from 19 publications to train and operate its models. The 79-page complaint, brought by the parent company alongside 13 affiliated entities, seeks more than $250 million in damages, a permanent injunction, and an order to destroy GPT models and training sets containing the plaintiffs' content.

The statutory math explains the headline number. The plaintiffs say they may recover up to $150,000 for each willfully infringed work, plus up to $25,000 for each time OpenAI stripped copyright management information, and the complaint alleges OpenAI's models copied the papers' reporting "on a massive scale" during training. The filing puts specific numbers on that: the papers account for more than 160,000 entries in WebText, the dataset OpenAI built to train GPT-2, and more than 122 million tokens in a 2019 snapshot of Common Crawl known as C4 - including 23 million tokens from usatoday.com alone.

What separates this suit from earlier publisher complaints is its focus on output, not just training. The papers' lawyers ran the same prompt against GPT-5.6 once per headline - "Please find and summarize the article with this title and give me an in-depth summary" - and walked through 19 numbered examples, from the Indianapolis Star to the Detroit Free Press, in which the model produced detailed summaries of the underlying reporting. A link next to a full summary does not bring readers back, the complaint argues; the summary is the substitute.

The plaintiffs also put OpenAI's own words in evidence. The complaint quotes ChatGPT chief Nick Turley writing that publishers face an "existential threat" from OpenAI's products and are "largely substitutive, period," cites an OpenAI engineer's remark that "no matter how prominently we show the links, users won't click," and references internal descriptions of ChatGPT as the "modern newsstand." The suit asserts claims for direct and vicarious copyright infringement plus removal of copyright management information under the DMCA, names seven OpenAI entities, and demands a jury trial.

The 19 publications span USA Today and papers including The Tennessean, Indy Star, the Detroit Free Press and The Detroit News, The Arizona Republic, The Columbus Dispatch, The Oklahoman, the Milwaukee Journal Sentinel and The Palm Beach Post. The plaintiffs filed a statement of relatedness the same day asking that the case be consolidated with the OpenAI copyright litigation already pending before the same court.

USA Today's publisher joins a crowded docket. The New York Times sued OpenAI and Microsoft in December 2023; since then The Intercept, Ziff Davis, CBC/Radio-Canada, Encyclopaedia Britannica, Merriam-Webster, The Seattle Times and a coalition of nearly 400 local newspapers have filed claims of their own, many of them consolidated in the same Manhattan courthouse.

OpenAI had not publicly responded to the suit as of publication. The company has signed paid licensing deals with other news organizations - something the complaint itself cites as proof that OpenAI knows a license is required - while it fights the litigation on multiple fronts and prepares for an eventual public listing.

The practical stake goes beyond one publisher. If courts start treating these cases as a pattern rather than isolated disputes, settlements or licensing schemes could set a de facto price for news archives across the industry. A $250 million demand from a single mid-sized publisher is also a signal about how plaintiffs value their archives - and every AI company that trained on scraped news is watching the same numbers.

Comments (0)

Log in to join the discussion

Log In

No comments yet