AI firms threaten physical books, urging urgent digitisation of rare titles.
| Source: HN | Original article
AI firms are buying and destroying physical books to train models, prompting calls to scan rare books before they vanish.
AI firms are quietly buying millions of second‑hand books, digitising them and then destroying the physical copies, a practice exposed in a guest post on the volunteer‑run Anna’s Archive. The post alleges that the books – including rare and out‑of‑print titles – are scanned en masse to feed the massive language models that power today’s generative AI services.
The drive to acquire physical volumes appears to have begun after companies exhausted the readily available online text corpus. As Emmanuel Maiberg notes, the rush for rare books is a response to dwindling digital sources, prompting firms such as Anthropic and Amazon to purchase and then discard the originals. The Hindu’s coverage echoes the claim, reporting that AI firms have also downloaded pirated e‑books before turning to physical collections.
The implications are stark. By converting unique, often irreplaceable works into proprietary digital files, a handful of corporations could become the sole custodians of vast swaths of cultural heritage. Critics describe the practice as “a crime against humanity,” warning that knowledge would be permanently monopolised on private servers, inaccessible to scholars, libraries and the public.
In response, Anna’s Archive has launched a volunteer‑driven campaign to scan and preserve at‑risk books before they disappear. The initiative seeks to create an open, searchable repository that can counterbalance the private hoarding of digitised content. Observers will be watching for any regulatory moves aimed at protecting physical collections, as well as for broader industry reactions to the preservation effort. The coming weeks may see heightened debate over intellectual‑property law, cultural‑heritage safeguards and the ethical limits of data‑harvesting for AI.
Sources
Back to AIPULSEN