My analysis shows that Amazon is running a covert operation to mass-digitize printed publications, one that shocks even seasoned antiquarian booksellers. This is not just about scanning—the books are physically destroyed after processing. It is a systematic approach that reveals tech giants' appetite for unique data to train neural networks.
How the scheme came to light
It all started with anomalous orders from rare book sellers. Anonymous buyers were purchasing hundreds of editions of varying ages and topics, including unique copies. In July, one antiquarian bookstore received an order for a thousand books through the Biblio platform. The seller, suspecting something was off, agreed to track the fate of the batch by hiding an Apple AirTag GPS tracker inside one of the volumes.
The device led to an Amazon facility in North Las Vegas. There, the VGT3 division is located, which, as it turned out, handles the processing of printed books. By studying employee posts and speaking with former workers, I pieced together the full picture of what is happening.
A conveyor belt for destroying knowledge
The process looks like this. First, employees register each edition and scan barcodes or ISBNs. Then comes the most barbaric part: the book bindings are cut off, turning them into stacks of individual pages. These pages are loaded into high-speed scanners. This approach maximizes the speed of digitization but makes restoring the physical book impossible.
According to workers, scanning is VGT3's primary function. There is also indirect evidence of the scale: in early 2026, the platform even experienced supply disruptions—so great was the demand for new batches of books for processing.
Official position and reality
Amazon, commenting on the situation, confirmed that it purchases printed editions through commercial channels. The corporation stated that the books are used for "developing and improving products and services." However, the company stubbornly remains silent about the scale of purchases, the program's timeline, and the specific products for which the data is being prepared.
This silence speaks louder than any words. It is obvious that Amazon is building its own pool of training data, preferring physical archives over digital ones—likely due to issues with copyright and the quality of existing datasets.
My verdict: We are witnessing a new arms race in the AI industry. Buying and destroying rare books is not just a way to obtain text, but a strategic move to create exclusive content for training. While regulators argue about digital piracy, corporations are mastering far more aggressive data collection methods that remain outside the legal framework. The only question is when the public will realize the price we pay for the progress of algorithms.