The Pulse
Unsealed Filings Show OpenAI and Microsoft Saw News as a Threat
Unsealed court filings reveal internal OpenAI and Microsoft warnings that AI products could replace news publishers and damage the web’s supply of reporting. The documents also detail scraping plans, paywall circumvention, data sharing and

AI.info Team ·
“The largest theft of labor in human history.”
Brent Hecht, Microsoft director of applied science, as quoted in an unsealed court filing
Microsoft and OpenAI’s internal debate over news scraping became public on September 17, when a federal court unsealed portions of a summary-judgment motion filed by news organizations led by The New York Times. The documents cited in the motion describe executives and researchers who feared that AI products trained on journalism could undermine the publishers whose work supplied the systems.
The filings are part of the copyright case brought by The New York Times and other news organizations against OpenAI and Microsoft. They do not resolve whether the companies’ conduct qualifies as fair use. They do expose internal statements that the publishers say conflict with the companies’ courtroom position that model training is transformative and that products such as ChatGPT and Copilot do not replace news websites.
Microsoft’s “doom loop” warning
Hecht, Microsoft’s director of applied science, described the copying of news for AI training as “an astonishing theft of unprecedented proportions” and possibly “the largest theft of labor in human history,” according to the publishers’ motion. In another document, he wrote that treating the practice as fair use would make “a complete mockery of the idea of ‘fair use.’”
A separate Microsoft document warned of a “doom loop” in which AI products would weaken the economic base of the publishers supplying their training material. “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers,” the document said, “but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”
Microsoft’s internal data, as presented in the motion, showed click-through rates falling by 83% to 93% for some news plaintiffs and by 51% to 94% for others. The publishers argue that those figures support their claim that chatbot answers and AI search features can substitute for visits to the original sites.
OpenAI messages describe substitution
OpenAI’s internal messages carried a similar concern. Nick Turley, identified in the filing as the head of ChatGPT, wrote that publishers faced an “existential threat” from commercial products trained on their work. He described the products as “largely substitutive, period,” and predicted that they would become more substitutive as they improved.
Another OpenAI employee wrote that users would not click even when chatbots displayed prominent links. Turley agreed that there was “no good reason to click” when a chatbot supplied the requested information directly. Microsoft CEO Satya Nadella also testified that conversational AI had substituted for publisher websites by providing information inside the AI product instead of sending users to the original source.
OpenAI President Greg Brockman was told by researcher Nick Ryder that the company had found “a hack” to get around The New York Times’ paywall for its crawlers. Brockman replied, “Ah, nice,” according to the motion.
Project Mango and the scale of the copying
The filings describe data exchanges between the two companies. OpenAI provided Microsoft with the full GPT-3 training dataset so Microsoft could evaluate how to use OpenAI models in commercial products. Microsoft, in turn, supplied training data to OpenAI through projects called Project Taxi and Project Mango.
The publishers say Project Mango contained copies of at least 160,903 unique works from news organizations. They also allege that Microsoft sold or repurposed data gathered for Bing, while OpenAI obtained a third-party dataset containing 1.8 million New York Times articles despite restrictions against commercial use.
The motion further alleges that OpenAI built systems to strip copyright notices from training data because researchers did not want models producing those notices in their responses. The plaintiffs also say the companies created filters that suppressed output from publishers who had sued, while leaving content from other organizations less restricted. Hecht described that result as an “accidental cover up.”
Microsoft separates testimony from company policy
Microsoft disputes the publishers’ interpretation. A Microsoft spokesperson told Ars Technica that Nadella’s testimony concerned broad changes in how people find and consume information and should not be treated as a conclusion on the copyright issues before the court.
The company also said Hecht’s comments reflected “one employee’s individual perspective,” were not legal analysis and did not represent Microsoft’s views. Microsoft maintains that AI training qualifies as fair use and that Copilot does not substitute for publishers’ journalism. OpenAI did not respond to Ars Technica’s request for comment.
Steven Lieberman, counsel for the New York Daily News and seven affiliated papers, said the newly visible material showed that OpenAI and Microsoft knew their conduct was wrong. The companies have argued that the documents should remain confidential; the court’s unsealing order now puts the internal warnings into the public record.
The immediate legal question is whether the statements and traffic data will persuade the court that the companies’ use of news content harmed the market for the original work. The publishers are asking for judgment on a narrower group of articles where chatbot outputs allegedly show extensive verbatim overlap. The remaining disputes are headed toward the broader copyright fight over how OpenAI and Microsoft obtained, exchanged and used the material.