AI Developers Turn to Bulk Book Destruction to Feed Models and Avoid Legal Risk
Brokers buy pre-2023 print volumes to harvest human prose before pulping physical copies.
Artificial intelligence laboratories facing a shortage of clean training data and growing copyright liabilities are buying up massive quantities of physical print books, severing their bindings, scanning their pages, and throwing the physical remainders away.
By utilizing stealthy middleman procurement companies that guarantee client anonymity, tech firms are securing hundreds of thousands of print volumes at industrial scale. The surge in demand targets out-of-print literature, foreign-language titles, and obscure academic works, creating an unprecedented cash influx for second-hand booksellers while raising alarm among archivists who fear rare physical texts are being permanently lost to destruction.
The aggressive push toward physical print stems from two converging challenges in machine learning development: text quality and legal exposure.
As generative AI systems flood the internet with synthetic content, researchers are encountering data degradation—often referred to as model collapse—when training new systems on web content produced after late 2022. Books published before the public release of modern large language models offer a pristine repository of human thought, complex reasoning, and structured syntax unadulterated by machine-generated text.
Simultaneously, purchasing physical print offers AI developers a potential legal safe harbor. Under the first-sale doctrine in United States copyright law, the legitimate purchaser of a physical book holds the right to resell, alter, or destroy that specific physical copy. When paired with recent federal court interpretations of the fair use doctrine, converting those legally acquired physical pages into digital training data has gained significant legal protection.
This legal strategy was highlighted in the San Francisco federal court case Bartz v. Anthropic PBC, where a judge ruled that Anthropic’s “Project Panama”—a secret program that digitized millions of print books via destructive spine-cutting—constituted transformative fair use, even though the physical volumes were destroyed after scanning. Subsequent cases involving Meta Platforms and OpenAI yielded similar judicial reasoning regarding the transformative nature of training datasets.
That legal defense contrasts sharply with the liabilities associated with unauthorized digital repositories. In a separate proceeding in the same San Francisco federal court district, Anthropic agreed to a $1.5 billion settlement after admitting to using illicit digital “shadow libraries” to train its Claude models, resulting in payouts of approximately $3,000 per title to affected authors.
The shift to physical procurement has drastically altered the economics of used book dealing. Independent sellers report their weekly transaction volumes expanding exponentially, with obscure titles that sat unsold for years suddenly commanding immediate bulk purchases. While booksellers welcome the revenue, many express distress over the fate of their inventory, particularly when foreign and out-of-print works are systematically pulped after being scanned.
The destructive methodology has drawn criticism from within the technology sector as well. Following reports detailing the scale of destructive scanning, xAI founder Elon Musk publicly instructed his engineering teams at SpaceX and xAI to preserve physical volumes, directing them to build high-speed non-destructive scanning systems rather than slicing off book spines.








