
404 Media investigation tracks rare books to an Amazon facility where they are scanned and destroyed for AI training data.
AI-generated summary
AI companies seek pre-2022 printed books to avoid 'model collapse' caused by training on AI-generated content. Several major tech firms are currently involved in legal disputes regarding copyright and fair use in AI training.
AI companies are scanning truckloads of books, sometimes rare ones, and often destroying them in the process to train their models—and now we have some idea of which companies and how they’re doing it.
Independent tech media publication 404 Media placed a tracking device inside a shipment of rare books and watched it travel to an Amazon facility in Las Vegas where the company scans and destroys printed books to train AI.
The investigation, published Monday, is the first to publicly pinpoint where bulk book purchases tied to AI training end up.
The final stop was Amazon's VGT3 team. Employees told the outlet all they do is receive massive shipments of printed books, then cut the bindings off so the pages feed a scanner more quickly. The book is destroyed in the process. The team's extremely fitting logo is a dinosaur clutching a book in its claws.
Why printed books are suddenly in demand
Books printed before 2022 carry text that isn't mostly readily available online and, crucially, isn't machine-written. Training a model on its own kind of output risks "model collapse," where quality degrades with each recursive loop.
So, pre-2022 paper is clean fuel for AI training.
Amazon isn't the only one. A cottage industry has formed to supply AI labs with physical books that get stripped, scanned, and discarded—the "Fahrenheit 451" scene of texts destroyed to feed machines that Anthropic's internal "Project Panama" already ran at scale, digitizing millions of books through destructive scanning.
That scramble for clean text has already landed in court. A federal judge ruled that training AI on legally purchased books can qualify as fair use, a partial win for Anthropic that Meta and OpenAI also claimed. Pirated copies are a different matter: Anthropic agreed to a $1.5 billion settlement over scanned pirated titles, and Salesforce now faces a class action over alleged book piracy.
The Amazon operation stayed quiet for a structural reason. Booksellers noticed a spike in bulk orders over the past year and suspected AI companies were behind them—the buyers weren't price-sensitive and picked titles seemingly at random. Their concerns were finally confirmed by an AirTag revealing the destination.
Amazon said it "purchases books through commercial channels to help develop and improve the products and services our customers use."

MANTRA Chain halted its mainnet on Aug. 21 after an attacker exploited an upstream dependency. Transactions, staking, and transfers are currently suspended while the team tests a security patch on the DuKong testnet before a coordinated restart.

Solana has successfully reduced its slot time to 350 milliseconds, down from 400ms, as part of a multi-stage plan to improve network latency. The update, approved via SIMD-0525, aims for further reductions toward a 200ms target.

Ethereum's better.codes contest tracks a 52.14-bit cryptographic proof gap for the koalaIRS12 parameter profile, measuring distance between certified safety and unsafe bounds via soundness and attack tracks.

Coldcard maker Coinkite released a security overhaul for Bitcoin hardware wallets following a firmware flaw that led to over $130 million in stolen Bitcoin.

Solana has upgraded its network for the first time since genesis, reducing base slot timing from 400ms to 350ms to speed up transaction confirmations. The change is part of a phased plan to reach 200ms, aiming to improve latency and censorship resistance.

A Bitcoin address tied to Maya Protocol's Aug. 18 exploit still held ~20.8 BTC worth $1.59M on Aug. 21, as technical analyses reveal broader pool damage exceeding initial estimates and recovery plans remain undefined.