Microsoft director calls AI scraping “largest labor theft in history”
microsoft openai training
| Source: HN | Original article
A Microsoft director described AI data scraping as the largest theft of labor in human history, according to newly unsealed court filings.
A court brief filed by the New York Times and recently unsealed shows a Microsoft director describing the data‑scraping that fuels large language models as “the largest theft of labor in human history.” The internal comment, made by Brent Hecht, who runs applied‑science at Microsoft, appears in filings that also reveal the company and its partner OpenAI built training sets from pay‑walled Times articles and warned internally about the practice.
The revelation adds a stark contrast to the public stance of both firms, which have defended the use of web‑scale data as fair use. It also provides concrete evidence that senior executives were aware of the ethical and legal tensions surrounding the mass extraction of copyrighted content for AI training. The admission arrives amid a wave of litigation targeting the AI industry’s data practices and follows a series of disclosures this week about a “doom loop” in which AI models both rely on and erode the very web they consume.
Why it matters is twofold. First, the statement could bolster the New York Times’ case that the scraping infringes copyright and undermines the labor of journalists and other content creators. Second, it fuels broader policy debates in Europe and the United States about how to regulate AI training data, protect intellectual property, and ensure that the benefits of generative AI do not come at the expense of the creators whose work fuels it.
What to watch next are the outcomes of the Times’ lawsuit, potential regulatory responses, and whether Microsoft or OpenAI will adjust their data‑collection practices. Industry observers will also be looking for any further internal documents that could reveal how widespread the concern was among AI developers. As we reported on 19 September, the “doom loop” narrative is now backed by internal admissions that the scale of data extraction may indeed be unprecedented.
Sources
Back to AIPULSEN