Amazon found to be destroying rare books for AI training data
A 404 Media investigation reveals that Amazon is purchasing large quantities of rare books to scan for artificial intelligence, a process that reportedly destroys the physical copies.
A 404 Media investigation has revealed that Amazon is purchasing large quantities of rare books to scan for artificial intelligence training data, destroying the physical copies in the process. The report identified the operation by placing a tracking device in a rare book suspected of being acquired by an AI company and following it to its final destination. The shipment arrived at an Amazon warehouse in Las Vegas, Nevada, where the books were processed for digital extraction.
Employees at the Las Vegas facility, known internally as team VGT3, reportedly cut the bindings off printed books to facilitate rapid scanning. This method prioritises speed over preservation, resulting in the destruction of the physical items. The team’s logo is described by employees as a dinosaur brandishing its teeth and holding a book in its hands.
The tracked shipment contained rare books, some of which were in foreign languages or had limited print runs. While the specific titles have not been revealed, the books represent historical and intellectual value beyond their monetary worth. A bookseller who sold the volumes noted that AI companies primarily value the content as text, often disregarding the historical, intellectual, or sentimental value of the physical books.
The destructive method contrasts sharply with the approach of the Internet Archive, founded by Brewster Kahle in 1996. The Internet Archive digitises books page by page to preserve the physical integrity of the volumes, a practice that requires specialised machines and software. This method was chosen to avoid the trade-offs seen in earlier industrial digitisation efforts.
In 2004, Google’s Books project introduced industrial scanning methods that often involved cutting spines and dismantling bindings to prioritise speed. The Internet Archive explicitly rejected this approach, opting instead for a slower, more careful process that maintains the physical structure of the books.
The current controversy highlights a broader tension in the digital landscape, where Amazon claims “fair use” for its AI data acquisition. This stance stands in contrast to legal actions taken by large corporations against the Internet Archive over similar practices, raising questions about the standards applied to different entities in the race for training data.

