Why AI Companies Buy Then Destroy Old Books
google
| Source: Mastodon | Original article
AI firms are purchasing rare books only to later destroy them, sparking concerns over cultural preservation.
AI firms are quietly amassing millions of out‑of‑print volumes, digitising them and then discarding the originals. A series of reports this week – including a piece on Open Culture and coverage on Google News – describe how companies buy books “by the pallet”, slice them apart, scan the pages and pulp the remnants. The practice spans a bizarre range of titles, from Italian manuals on home‑oxygen treatment to medieval English marriage‑law treatises, modern Texas civil‑procedure guides and 1960s Swedish comedy collections.
The revelation matters because it pits the commercial drive to feed large‑language models with massive text corpora against the preservation of cultural heritage. Physical books, especially rare or out‑of‑print works, are often the only surviving copies of niche knowledge. Their destruction eliminates any chance of future scholarly access, even as the same content reappears in proprietary AI datasets that are not publicly searchable. Former U.S. House Representative Brad Carson highlighted the paradox on X, noting that “AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what’s left – here’s the perverse part.”
The story is likely to trigger scrutiny from both regulators and preservation groups. Watch for possible legislative or export‑control measures that could extend the U.S. rule‑making discussed earlier this month, aimed at curbing the flow of AI‑related materials. Industry bodies may also introduce transparency standards for training‑data sourcing, while libraries and cultural institutions could lobby for stricter protections against bulk purchases. The next weeks should reveal whether the practice prompts policy action or remains a hidden facet of the AI data‑harvest boom.
Sources
Back to AIPULSEN