Newly released court documents reveal significant internal concern at Microsoft and its close partner OpenAI over the use of millions of news articles to develop artificial intelligence systems.
As OpenAI continued expanding its AI technology, Microsoft employees discussed whether the company’s approach to using published material could amount to what one internal discussion described as the “largest theft of labor in human history.” Employees also raised concerns about the possibility of creating a “doom loop” in which AI systems could eventually undermine the very sources of high-quality information needed to train large language models.
According to a 2023 internal Microsoft document, employees anticipated that millions of people around the world could come to regard large AI models collecting and using their work as an unprecedented form of appropriation.
Similar concerns were being discussed within OpenAI. Executives considered the possibility that ChatGPT could reduce the number of people visiting news publishers directly by providing information within the chatbot itself.
In a June 2023 internal memo, Nick Turley, who led the team responsible for developing ChatGPT, described artificial intelligence as posing an “existential threat” to publishers. By February 2024, Turley wrote that AI products were likely to become increasingly capable of substituting for existing sources as the technology improved.
Parts of these internal discussions became public on Thursday through court filings connected to the copyright lawsuit brought by The New York Times against OpenAI and Microsoft in late 2023. Eleven other publishers have since joined the litigation.
The case is being considered by Judge Sidney H. Stein of the U.S. District Court for the Southern District of New York. He is currently reviewing motions seeking summary judgment, while additional documents connected to the litigation are gradually being unsealed.
The newly disclosed material provides a closer look at concerns that existed inside the companies themselves as generative AI systems rapidly developed, particularly over how the technology could affect publishers, journalism and the production of original material.









