AI Companies Buy, Scan, and Destroy Books: Pulping Rare Volumes Sparks Cultural Preservation Debate

AI companies are buying, scanning, and pulping books—including rare editions—for training data.
Some AI companies are industrially purchasing physical books, destructively scanning them for training data, and then pulping the remains—including rare antiquarian and out-of-print volumes. While legally defensible under the first-sale doctrine, this practice raises serious cultural preservation concerns as irreplaceable physical artifacts are permanently destroyed for commercial AI development.
Overview: Physical Books Are Being Destroyed at Scale for AI Training
A Reddit discussion is drawing widespread attention across the tech community: to train large AI models, some AI companies are systematically purchasing, scanning, and physically destroying paper books — including rare antiquarian volumes and out-of-print editions.
According to the original reports, this process has become highly industrialized. Companies use hydraulic cutting machines to slice off book spines, feed the loose pages into industrial-grade scanning equipment for bulk digitization, and then pipe the extracted text into AI training systems. Once the process is complete, the dismantled physical books are typically treated as waste paper and sent straight to pulp recycling.

Notably, this isn't an isolated practice by a single company — it's becoming an industry-wide norm. Used book dealers have even begun treating the AI boom as a business opportunity, proactively buying up secondhand books to resell to AI teams hungry for training data. A gray-market supply chain built around "book digitization and destruction" has taken shape.
Why Do AI Companies Choose to "Destroy" Rather Than "Preserve" Books?
From an engineering perspective, the core motivation for destroying books is scanning efficiency. Traditional non-destructive scanning requires manual page-turning, flattening each page against the spine — it's slow, expensive, and performs poorly on thick or tightly bound volumes. Destructive scanning, which involves removing the spine and separating the pages, allows the use of automatic document feeders for high-speed processing, improving efficiency by orders of magnitude.
It's worth diving into the technical details of destructive scanning. Destructive scanning is a document digitization method that sacrifices the integrity of the original in exchange for extremely high processing speed. The core equipment is a high-speed Automatic Document Feeder (ADF), paired with hydraulic cutters or rotary trimming devices to remove the spine. Once separated, the loose pages can be fed through one by one — just like copy paper — and scanned on both sides simultaneously. Industrial-grade devices like the Fujitsu fi series or Kodak i5000 series can process over 200 pages per minute, with OCR (Optical Character Recognition) software converting images to searchable text in real time. By comparison, non-destructive scanning requires V-shaped book cradles or planetary scanners, with operators manually turning pages and adjusting focus — each page can take upwards of 30 seconds. This efficiency gap — from 120 pages per hour to 12,000 pages per hour — is the direct reason AI companies opt for the destructive approach.
For AI companies chasing massive volumes of high-quality training corpora, physical books represent an extremely valuable category of training data: they've been professionally edited, fact-checked, written in standardized language with high information density, and are largely free from the quality issues that plague repeatedly crawled web data. Compared to noisy web text, book-sourced corpora significantly improve a model's knowledge depth and expressive quality.
To understand why books are so important for AI training, you need to understand the quality hierarchy of training corpora. In the training pipeline of large language models (LLMs), data quality directly determines the model's final performance. The industry generally classifies training corpora into several quality tiers: at the bottom is uncleaned web crawl data (such as Common Crawl), which is enormous in volume (petabyte-scale) but riddled with ad copy, SEO spam, and duplicate content; the middle tier includes Wikipedia, news sites, and other texts that have undergone some editorial process; the top tier consists of professionally edited and fact-checked publications — namely books and academic papers. Research has shown that, at equivalent data volumes, models trained on high-quality book corpora significantly outperform those trained purely on web data in terms of factual accuracy, logical coherence, and linguistic quality. OpenAI's GPT series, Google's PaLM, and other models have all heavily utilized book datasets like Books Corpus in their training data. As the Chinchilla scaling law revealed the optimal ratio between data volume and model parameters, the industry's demand for high-quality data has become even more urgent.
The Legal "Compliance" Gray Zone
Reports specifically note that this practice occupies a legally defensible space under U.S. law, primarily relying on two principles:
- First-sale doctrine: Once you've legally purchased a physical copy of a book, you have the right to dispose of that copy — including reselling or destroying it — and the copyright holder cannot interfere with the fate of that physical object.
- Fair use: Under certain conditions, transformative use of a work for purposes such as analysis or research may qualify as fair use.
The first-sale doctrine has deep legal roots. It originates from the U.S. Supreme Court's 1908 decision in Bobbs-Merrill Co. v. Straus and was later codified in Section 109 of the U.S. Copyright Act. Its core meaning is that a copyright holder's control over a specific copy ends once that copy has been lawfully sold for the first time. The buyer is free to resell, lend, rent (with exceptions for music recordings and software), display, or even destroy the copy without the copyright holder's consent. This principle is the legal foundation that makes secondhand bookstores, library lending systems, and thrift markets possible. However, the doctrine only covers the right to dispose of the "physical copy" — it does not extend to copying or reusing the copyrighted content within it.
Fair use determination is far more complex, requiring a rigorous four-factor test. Section 107 of the U.S. Copyright Act requires consideration of: (1) the purpose and character of the use, particularly whether it is "transformative" — creating new meaning or value rather than simply substituting for the original; (2) the nature of the copyrighted work, with factual works receiving weaker protection than creative works; (3) the amount and substantiality of the portion used, where full-text copying generally weighs against fair use claims; and (4) the effect on the potential market for the original. The key ongoing lawsuits around AI training — such as Authors Guild v. OpenAI, Sarah Silverman v. Meta, and The New York Times v. OpenAI — are all fiercely debating how these four factors apply in the AI context. AI companies argue that training is a highly transformative use because models generate statistical patterns rather than reproducing original text; copyright holders counter that entire books are "ingested" into training systems and that model outputs directly compete with original works, constituting substantive infringement.
In other words, when AI companies buy books, scan them, and destroy them, they're difficult to hold liable at the level of "disposing of physical copies they own." But this is only legality at the physical copy level. Whether using scanned content for model training constitutes infringement remains the focal point of multiple ongoing AI copyright lawsuits — and is far from settled.
The Core Controversy: What Is the Cultural Cost of AI Progress?
What has truly ignited heated community debate isn't the question of legality, but the cultural and ethical cost.
When only commonly available copies of bestsellers are being destroyed, the controversy is relatively limited — after all, countless identical copies still exist. But when the targets expand to rare books, out-of-print editions, and antiquarian volumes with only a handful of surviving copies, the situation is entirely different. Each copy of such a book may be an irreplaceable cultural artifact. Once it's cut apart and pulped, its value as a historical artifact, a work of bookbinding art, and a vessel of marginalia and annotations is permanently lost.
In library science and conservation studies, a book's value extends far beyond the text it carries. Materiality is a key concept for understanding the cultural heritage value of books. The material, thickness, and chemical composition of the paper can help date the work and trace its origin; traces of movable type or hand typesetting record the evolution of printing technology; bookplates, seal impressions, and annotations form a complete "provenance chain" reflecting the book's path through different eras and readers. For example, handwritten notes in the margins of an 18th-century volume might document a scholar's thought process, with historical research value potentially exceeding that of the printed text itself. Furthermore, a book's binding — leather covers, gold-stamped tooling, hand-sewn signatures — constitutes physical evidence of craft and decorative arts. All of these dimensions are completely erased in any form of text-extraction scanning. Even high-resolution image scanning cannot fully preserve tactile qualities, scent, and three-dimensional structural information.
Digitization Is Not "Preservation": The Overlooked Information Loss
A common defense is that the content has been scanned and preserved as digital text, so the disappearance of the physical object doesn't mean knowledge has been lost. But this argument overlooks several critical facts:
- Information loss: Scanning typically extracts only text content. A book's physical characteristics — paper quality, printing techniques, handwritten marginalia, ownership stamps, binding — all the dimensions that carry historical information are completely discarded.
- Privatization risk: The digital copies of destroyed books are held by AI companies and used for commercial model training. The public may never have access to them — a fundamentally different paradigm from public preservation in libraries and archives.
- Irreversibility: Unlike high-precision conservation scanning of endangered documents, destructive scanning prioritizes efficiency over preservation. The result is the complete and permanent disappearance of the original.
A Deeper Industry Concern: The Clash Between Data Hunger and Cultural Heritage
This phenomenon reflects a sharp contradiction at the heart of the current AI arms race: the insatiable demand for training data is coming into direct conflict with cultural heritage preservation.
As high-quality web text is gradually exhausted by crawling, models are increasingly desperate for "uncontaminated" premium corpora — and books have become one of the most tempting targets. The "data wall" problem facing the AI industry today is more severe than the public realizes. According to estimates by the Epoch AI research institute in 2022, the total volume of high-quality text data available on the internet for training is on the order of several trillion tokens, and the latest generation of large models (such as GPT-4, Llama 3, etc.) have already consumed a significant portion. At the current rate of model scale growth (scaling laws require data volume to expand in tandem with parameter count), high-quality text data could be effectively "exhausted" around 2026. This is forcing AI companies to continually explore new data sources: synthetic data generation, multimodal-to-text conversion, and — as discussed in this article — converting not-yet-digitized physical books into training corpora. It's estimated that approximately 130 million unique published books exist worldwide (according to the Google Books project), a large proportion of which have never been digitized — especially out-of-print editions and publications in less common languages. This vast "unmined" corpus presents an almost irresistible temptation for data-hungry AI companies.
In the absence of clear industry standards and regulation, efficiency and cost will naturally override careful consideration of cultural value.
There's a warning signal worth noting here: when destroying rare books becomes a business that can be "monetized," market forces will accelerate the process. The profit-seeking behavior of used book dealers, the data hunger of AI companies, and the ambiguity of the law — the convergence of these three factors could produce irreversible cultural losses, the kind that are often only truly recognized years later.
Conclusion: We Need Clear Boundaries Between AI Training and Cultural Preservation
We must acknowledge that destructive scanning has its rationale in terms of technology and business logic, and for commonly available books, it's not necessarily unacceptable. The real problem lies in the lack of boundaries — should protective red lines be established for rare and out-of-print books? Should non-destructive scanning and public archiving be prioritized before any destructive digitization? Does the industry need to establish self-regulatory standards?
This debate over antiquarian books and AI training is fundamentally asking: how great an irreversible cultural cost are we willing to pay for AI "progress"? In an era of breakneck technological advancement, this question deserves serious consideration from every practitioner and user.
Note: This article is based on reports circulating in the Reddit community. The specific scale of these practices and the companies involved have yet to be verified by additional authoritative sources. Readers should exercise critical judgment.
Related articles

Getting Started in Machine Learning Research: Essential Paper Reading List and Research Internship Application Path
A complete path from zero to research internship for ML beginners, covering essential classic papers (AlexNet, ResNet, Transformer), paper reading methods, reproduction tips, and practical advice for research internship applications.

Claude Code Hands-On Tutorial: Complete Guide from Installation to Automated Development
Complete guide to Claude Code covering environment setup, permission configuration, Go Goals autonomous loops, Skills system, MCP protocol integration, and version control for automated development.

Gemini 3.7 Flash Release and GPT-5.6 Ultra-Fast Mode: AI Open Source Enters the Ecosystem Era
Google releases Gemini 3.7 Flash for coding and Agent optimization while OpenAI launches GPT-5.6 Ultra-Fast mode with 14x speed gains. AI open source shifts from open models to open ecosystems.