Inside the Warehouse Where Amazon Scans and Destroys Books for AI Training
amazon training
| Source: Mastodon | Original article
Amazon operates a warehouse where it scans and destroys thousands of books to create training data for its AI systems, according to an employee interview.
Amazon’s VGT3 warehouse in Las Vegas has become the focus of renewed scrutiny after a new interview with an employee revealed how the company turns physical books into AI training data. The staff member, who asked to remain anonymous because they are not authorized to speak publicly, described a workflow that begins with bulk purchases of titles, followed by high‑speed scanning and the systematic destruction of the originals once digitised. The operation, which the employee says handles “thousands of books,” is housed in the same facility identified last week by an Apple AirTag trace that linked a bulk shipment of about 1,000 titles to Amazon’s LAS8 site, also known as VGT3.
The revelation builds on our earlier report on Aug 24, which documented the tracking of a rare book to the same Amazon facility and highlighted the broader practice of scanning and discarding physical volumes for AI model training. The interview adds a human perspective to the process, confirming that the destruction is intentional and not an accidental by‑product of scanning.
Why it matters is twofold. First, the practice raises fresh copyright and intellectual‑property questions, especially as some of the scanned works include rare or out‑of‑print titles that may still be under protection. Second, the lack of transparency about data provenance could affect the credibility of AI systems that rely on these texts, prompting concerns from authors, publishers and regulators about consent and compensation.
Going forward, observers will watch for Amazon’s response—whether the company will adjust its sourcing policies, provide more disclosure, or face regulatory action. Industry analysts are also tracking how other AI developers might react, potentially tightening their own data‑collection practices or seeking alternative, licensed sources. The story underscores a growing tension between the rapid expansion of AI training pipelines and the need for responsible, rights‑respecting data handling.
Sources
Back to AIPULSEN