Amazon confirms rare book acquisition for AI training data
Investigation by 404 Media reveals the tech giant is purchasing out-of-print volumes to feed large language models, citing the need for pre-2022 data unavailable online.

Amazon has confirmed it is purchasing rare books, removing their spines, and scanning the pages to train large language models. The practice was verified after 404 Media placed a tracking device in a rare book that was delivered to Amazon’s Las Vegas facility, known as VGT3. The company stated it acquires books through commercial channels to enhance customer products and services.
The VGT3 facility, which identifies itself with a symbol of a dinosaur holding a book in its claws, serves as a key node in this data acquisition strategy. Amazon told 404 Media in a statement that it purchases books through commercial channels to improve the products and services customers use. This admission follows an investigation that traced the physical journey of a tracked volume to the facility.
Large language models require vast amounts of text data for training, and companies like Amazon have already ingested available online content. Rare texts are prioritised as training data because they predate 2022 and are not available online. This specific timeframe is critical, as texts published before 2022 are considered valuable because they are unlikely to have been written by an LLM.
The focus on pre-2022 publications is a direct response to the risk of model collapse. This technical term refers to the degradation of output quality when AI models ingest too much AI-generated content. By sourcing human-written texts that are out of print or impossible to find on the internet, Amazon aims to maintain the integrity of its model outputs.
While the primary report focuses on Amazon’s operations, the source material notes that Anthropic has previously ingested illegally pirated books. The acquisition of physical rare books represents a shift in how major technology firms are securing the high-quality, human-generated data necessary to sustain the development of advanced artificial intelligence systems.

