Authors Say ChatGPT Built on Hidden Mass Piracy
openai
| Source: HN | Original article
Authors allege that OpenAI built ChatGPT using copyrighted material obtained without permission, describing the practice as concealed mass piracy.
OpenAI’s flagship model, ChatGPT, is now at the centre of a copyright dispute after a group of authors testified that the service was trained on “concealed mass piracy.” In a court filing that is heavily redacted, the plaintiffs allege that OpenAI sourced large numbers of books from the illicit library LibGen, downloading torrented copies rather than purchasing the works. The authors claim the company never bought the texts it used to train the language model, effectively building the product on stolen material.
The allegation matters because it strikes at the core of how generative‑AI systems acquire the data that powers them. If a court finds that OpenAI relied on pirated books, the ruling could set a precedent for how publishers and authors enforce copyright against AI developers. It would also intensify pressure on the industry to disclose training‑data provenance and to adopt more transparent licensing practices. The case adds to a growing wave of legal challenges that have already seen authors and publishers push back against AI firms over data use, as reported in our earlier coverage of the Anthropic settlement dispute.
What to watch next is the court’s handling of the redacted filing and whether the parties move toward a settlement. A definitive judgment could trigger broader litigation across the sector, prompting AI companies to audit their data pipelines or negotiate blanket licences with rights holders. Regulators in Europe and the United States are also monitoring the issue, and any legislative response could reshape the balance between innovation and intellectual‑property protection in the AI era.
Sources
Back to AIPULSEN