Amazon Warehouse Scans and Destroys Thousands of Books for AI Training Data, Employee Reveals

Anonymous worker describes cutting spines, scanning pages, and discarding remains at Las Vegas facility as scrutiny over AI training data practices intensifies

edit
By LineZotpaper
Published
Read Time3 min
An Amazon employee at the company’s VGT3 warehouse in Las Vegas has confirmed that the facility is used to destructively scan thousands of books—including rare titles, library discards, and UK government documents—for use as AI training data, raising fresh questions about copyright compliance and the ethical sourcing of machine learning datasets.

The revelation comes from an anonymous interview published by 404 Media, following an earlier investigation that tracked a shipment of rare books to the same warehouse. The employee described walking past rows of scanners and seeing colleagues cut the spines off books before feeding loose pages through high-speed scanners. The scanned pages are then discarded into large open boxes, “all mixed together,” making reconstruction impossible.

“When we first started a lot of them were brand new. Some of them are used as well,” the employee said. “We were even getting boxes of stuff from London. There were a lot of books from the University of London. There was even … kind of stapled together papers that said that they were presented to Parliament on the behalf of Her Royal Majesty the Queen.”

The employee noted that workers were not explicitly told the purpose of the operation, but many suspected the data was being used for AI training. “I think people who work there probably knew what was going on over there, but probably just didn’t want to talk about it,” they said.

The VGT3 warehouse shares a facility with LAS8, Amazon’s print-on-demand operation, both part of a larger Amazon complex in Las Vegas. The earlier 404 Media investigation placed a tracking device in a shipment of books a bookseller believed was being acquired by an anonymous AI company; the shipment ended at VGT3.

Amazon has not publicly commented on the specific operation. The company has previously stated that it respects copyright and does not use customer content for AI training unless explicitly permitted. However, the scale and secrecy of the book-scanning operation—and the destruction of physical books—may intensify scrutiny from authors, publishers, and regulators who argue that training AI on copyrighted works without compensation constitutes infringement.

The practice is part of a broader trend: AI companies, including Google, Meta, and OpenAI, have been sued for training models on copyrighted books, articles, and other works. The use of Amazon’s logistics network to acquire and process physical books suggests an effort to obtain high-quality, out-of-print, or hard-to-find texts that may not be available in digital form.

Lawmakers in several countries have proposed legislation requiring transparency in AI training data sourcing. The European Union’s AI Act mandates disclosure of copyrighted material used in training, while the US Copyright Office has launched inquiries into the issue.

§

Analysis

Why This Matters

  • The revelation that Amazon is operating a physical book destruction pipeline for AI training data exposes a hidden dimension of the AI data race, where even rare and archival materials are being consumed.
  • For authors and publishers, it suggests AI companies—potentially including Amazon itself—may be bypassing licensing deals and destroying physical copies that could otherwise be sold or preserved.
  • The lack of transparency raises regulatory questions: if companies can acquire and destroy copyrighted books without oversight, existing copyright protections may be effectively unenforceable.

Background

The practice of training AI on copyrighted works has been controversial since at least 2023, when authors including Sarah Silverman and George R.R. Martin filed lawsuits against OpenAI and Meta. In 2024, The New York Times sued OpenAI for using its articles. Amazon itself has developed AI models for its Alexa, retail, and AWS businesses, but has not disclosed its training data sources. The VGT3 warehouse came to light after a bookseller became suspicious of bulk purchases by an anonymous buyer and planted a tracker. Last week, 404 Media first identified the facility as the destination. This latest employee account provides firsthand corroboration of the destructive scanning process.

Key Perspectives

Amazon: The company has not commented on VGT3 specifically. In general, Amazon states it respects intellectual property and does not train AI on customer data without consent. It may argue that the books are either public domain, licensed, or that the scanning falls under fair use for research purposes. Authors and Publishers: Trade groups like the Authors Guild and the Association of American Publishers view this as mass copyright infringement. The destruction of physical books—effectively removing them from circulation—compounds the harm, as it prevents future sales or library access. AI Companies and Data Gatherers: The acquisition of diverse, high-quality text is essential for training large language models. Physical books offer a trove of data not easily found on the web. However, the secrecy and scale suggest a recognition that public disclosure could spark legal challenges.

What to Watch

  • Legal responses: Authors or publishers may file a class-action suit if they can identify specific books destroyed. Watch for subpoenas or discovery motions targeting Amazon.
  • Amazon’s official statement: The company may clarify the purpose of VGT3—whether for internal AI development, sale to third parties, or a different use.
  • Regulatory action: The EU or US Copyright Office may open an investigation. The involvement of British government documents could also draw UK parliamentary scrutiny.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.