How the investigation worked
The investigation by 404 Media reached its conclusion through an unusual but effective method: an AirTag was placed inside a shipment of rare books, and the box was tracked across the country.
The shipment ended up at an Amazon warehouse in Las Vegas, home to a team called VGT3, whose logo features a Tyrannosaurus rex holding a book. Workers there say they cut off book spines to speed up scanning, destroying the copies in the process. Amazon uses the scanned data to train its Nova models. A spokesperson said the company buys books through commercial channels to improve its products.
Why printed books are so valuable
Booksellers suspect that AI companies are trying to systematically scan every book by ISBN number. There are two concrete reasons printed texts are so sought after:
- Many of these texts do not exist online at all, so they cannot be obtained by crawling the web.
- Because they predate 2022, they are free of AI-generated content. Since training models on their own output is known to degrade quality, human-written text from before that point keeps growing in value.
Those two properties turn an ordinary book sitting on a library shelf into an irreplaceable resource for model training.
Amazon is not alone
Amazon is not the only company running this practice. A lawsuit by book authors revealed Anthropic's operation called "Project Panama": the company bought books on marketplaces, cut off their spines and digitised them.
The judge in that case ruled the scanning qualified as fair use and did not violate copyright. One part of the reasoning stands out: because the printed originals were destroyed, the copies were not reproduced and resold. The destruction of the book, in other words, counted in the company's favour in the legal assessment.
The heart of the debate
What makes the practice controversial is not only copyright. Both companies are turning sometimes rare books into private training material, and doing so by destroying the originals. Copies that may be irreplaceable disappear; knowledge is pulled off public shelves and locked inside the closed model of a single corporation.
The asymmetry here is plain. Once a book is physically gone, access to that text runs only through the product of the company that scanned it. Model output neither cites the source nor returns the text itself. The court's fair use assessment may have closed the legal question for now, but the question of cultural record remains: is a library still a library once it is stored inside the model that read it?