‘Doom Loop’: OpenAI and Microsoft Admit LLMs Is Destroying the Web and Built on Theft
microsoft openai
| Source: Mastodon | Original article
OpenAI and Microsoft acknowledge that large language models are harming the web and rely on stolen data, raising legal and ethical concerns.
Executives at Microsoft and OpenAI have publicly acknowledged what critics have long alleged: the large‑language models (LLMs) powering ChatGPT and other generative‑AI products are built on a massive, systematic appropriation of online content. Unredacted court filings in the New York Times v. OpenAI copyright lawsuit reveal a Microsoft executive describing the practice as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” An internal Microsoft memorandum goes further, warning that the company’s AI content strategy has created a “doom loop” that simultaneously degrades model performance and erodes the health of the open web.
The admission matters because it ties the technical architecture of today’s LLMs directly to legal risk. By scraping billions of webpages, news articles and other copyrighted material without permission, the models not only train on that data but also generate outputs that compete with the original sources for clicks and ad revenue. If courts accept the argument that this constitutes unlawful copying, the fallout could reshape the economics of AI development, force a redesign of data‑collection pipelines, and trigger sweeping regulatory scrutiny across the sector.
The next weeks will focus on the outcome of the New York Times lawsuit and any related motions that could set precedent for how AI‑trained content is treated under copyright law. Stakeholders will also watch for responses from Microsoft and OpenAI—whether they propose licensing frameworks, alter training data practices, or lobby for legislative carve‑outs. Parallel to the legal battle, policymakers in the EU and the United States are expected to intensify hearings on AI‑driven content theft, making the “doom loop” narrative a focal point for forthcoming regulatory debates.
Sources
Back to AIPULSEN