The revelation comes from an anonymous interview published by 404 Media, following an earlier investigation that tracked a shipment of rare books to the same warehouse. The employee described walking past rows of scanners and seeing colleagues cut the spines off books before feeding loose pages through high-speed scanners. The scanned pages are then discarded into large open boxes, “all mixed together,” making reconstruction impossible.
“When we first started a lot of them were brand new. Some of them are used as well,” the employee said. “We were even getting boxes of stuff from London. There were a lot of books from the University of London. There was even … kind of stapled together papers that said that they were presented to Parliament on the behalf of Her Royal Majesty the Queen.”
The employee noted that workers were not explicitly told the purpose of the operation, but many suspected the data was being used for AI training. “I think people who work there probably knew what was going on over there, but probably just didn’t want to talk about it,” they said.
The VGT3 warehouse shares a facility with LAS8, Amazon’s print-on-demand operation, both part of a larger Amazon complex in Las Vegas. The earlier 404 Media investigation placed a tracking device in a shipment of books a bookseller believed was being acquired by an anonymous AI company; the shipment ended at VGT3.
Amazon has not publicly commented on the specific operation. The company has previously stated that it respects copyright and does not use customer content for AI training unless explicitly permitted. However, the scale and secrecy of the book-scanning operation—and the destruction of physical books—may intensify scrutiny from authors, publishers, and regulators who argue that training AI on copyrighted works without compensation constitutes infringement.
The practice is part of a broader trend: AI companies, including Google, Meta, and OpenAI, have been sued for training models on copyrighted books, articles, and other works. The use of Amazon’s logistics network to acquire and process physical books suggests an effort to obtain high-quality, out-of-print, or hard-to-find texts that may not be available in digital form.
Lawmakers in several countries have proposed legislation requiring transparency in AI training data sourcing. The European Union’s AI Act mandates disclosure of copyrighted material used in training, while the US Copyright Office has launched inquiries into the issue.