AI Companies Are Destroying Rare Books to Obtain Training Data

AI companies are destroying rare books through destructive scanning to feed the growing demand for training data.
To satisfy the insatiable demand for high-quality AI training data, some companies are using destructive scanning to shred physical books at scale—including rare and irreplaceable editions. This article explores the efficiency logic behind destructive scanning, the unique value of book text as training data, the irreversible cultural heritage losses, and the copyright gray areas, raising urgent questions about the hidden costs of AI's data hunger.
A Quiet Dismantling of Knowledge
Recently, a discussion on Hacker News titled "AI companies are shredding rare books" sparked widespread attention in the tech community, quickly garnering 88 upvotes and 34 comments. The discussion revealed a little-known yet highly controversial phenomenon: to obtain high-quality training data, some AI companies and data providers are dismantling, scanning, and even destroying physical books at scale—including rare editions.

Behind this phenomenon lies the enormous appetite of current large language models for high-quality text data. Today's leading LLMs (such as GPT-4, Claude, Llama, etc.) are trained on datasets measured in trillions of tokens. OpenAI's GPT-3 was trained on approximately 300 billion tokens in 2020, and by 2024, frontier models' training data had grown by more than an order of magnitude. This exponential growth in data demand stems from research on Scaling Laws—the finding that model performance follows a power-law relationship with training data volume, and adding more data almost always yields predictable performance improvements. However, the total amount of high-quality English text available on the internet is finite. The research institution Epoch AI estimates that high-quality language data may be exhausted around 2026. As publicly available internet text gets scraped dry and faces mounting copyright lawsuits and compliance pressures, the professionally edited and proofread high-quality content contained in physical books is becoming a new target for AI training data.
Why "Destroy" Books to Scan Them: The Efficiency and Cost of Destructive Scanning
How Destructive Scanning Works
For bulk book digitization, the fastest and most economical method is ironically "destructive scanning." The specific process involves cutting off the book's binding, separating pages into individual sheets, and then feeding them through a high-speed document scanner for digitization. Industrial-grade equipment such as the Fujitsu fi series or Kodak i5000 series high-speed scanners, equipped with automatic document feeders (ADF), can scan 100-200 pages per minute. The entire workflow typically includes: using an industrial guillotine cutter to remove approximately 3-5 millimeters of the spine to release the pages, then placing the loose sheets into the scanner's feed tray.
By contrast, the "non-destructive" approach of scanning page-by-page with a flatbed scanner is much slower and far more labor-intensive. Non-destructive scanning uses V-shaped book cradles and high-resolution cameras (such as the Scribe system used by the Internet Archive), requiring manual page-turning and focusing for each page, achieving only 1/10 to 1/20 the efficiency of destructive scanning. The Internet Archive's Open Library project primarily uses non-destructive methods in its digitization process, but the disadvantages in speed and cost make it difficult to match the data acquisition scale of commercial companies.
In AI data acquisition scenarios that prioritize scale and speed, destructive scanning has become the default option. The entire process of cutting, scanning, and discarding a book can be highly automated, compressing per-unit costs to extremely low levels. When the target is millions of books, this efficiency advantage becomes particularly compelling.
The Scarce Value of Book Text Data
Book text holds a special position in AI training corpora. Compared to the low-quality, repetitive, and even AI-generated content that floods the internet, published books have undergone professional writing, editing, and fact-checking. They feature standardized language, coherent logic, and high knowledge density. This type of "clean" long-form text is crucial for improving LLMs' language capabilities and knowledge accuracy.
As high-quality public data becomes increasingly depleted, the industry has even raised concerns about a "data wall"—the worry that available high-quality training data may be exhausted within the next few years. Against this backdrop, books in libraries and used bookstores have become an untapped "data goldmine."
AI training data has formed a complex supply chain ecosystem. Upstream are specialized data brokers and data annotation companies (such as Scale AI, Appen, etc.); in the middle layer are data cleaning and quality control services; downstream are the major AI labs. In this chain, book scanning is typically performed by professional document digitization service providers—companies like 1DollarScan offer per-page scanning services. Some reports indicate that data brokers purchase physical books in bulk from used bookstores, deaccessioned library collections, and estate auctions, then digitize them in batches and sell the packaged text data to AI companies. This multi-layered outsourcing structure allows AI companies to maintain distance from the actual book destruction, while also creating a systematic absence of artifact value assessment.
The Core Controversy: The Cultural Heritage Cost of Destroying Rare Books
Irreversible Loss of Cultural Artifacts
What triggers the strongest emotional response is the fact that "rare books" are being destroyed. When ordinary, mass-printed modern books are dismantled, their content can still be obtained through other channels, and the physical loss is relatively limited. But when rare books, out-of-print editions, limited editions, or historically significant printed works are cut up and destroyed, their value as physical artifacts is permanently lost.
From the perspectives of bibliography and library science, the value of rare books extends far beyond their textual content. A book as a physical medium carries multiple layers of information: the material and watermarks of the paper can reveal its origin and era; binding craftsmanship (such as hand-stitching and gilt edges) reflects the publishing technology of a specific period; annotations, bookplates, and inscriptions within the book constitute a unique provenance history; even the wear patterns on pages can tell researchers which chapters were frequently read. According to the standards of the Antiquarian Booksellers' Association of America (ABAA), determining "rarity" considers factors including: surviving copies, historical importance, physical condition, and market demand. Once the physical object is destroyed, information across all these dimensions is permanently lost—a digital copy can only preserve text and images.
In the Hacker News comment section, many viewed this practice as a form of cultural shortsightedness—feeding irreplaceable cultural heritage to a commercial AI model. Digitization preserves textual content but cannot retain the historical information, binding craftsmanship, and collectible value that books carry as physical objects.
Divisions and Reflections in the Tech Community
The discussion also featured dissenting voices. Some tech professionals argued that for widely available ordinary books, destructive scanning in exchange for permanent digital preservation of content is actually a form of "rescue"—after all, paper books will decay over time, while digital copies can be infinitely reproduced and distributed. The real point of contention lies in "which books deserve physical preservation" and "who gets to make that judgment."
The true problem is the absence of a screening mechanism: in large-scale, automated data acquisition workflows, it's difficult to assess the artifact value of each individual book. When scanning work is outsourced to service providers focused on throughput, rare books and ordinary books are likely treated equally—both sent through the guillotine cutter.
Deeper Industry Concerns: Copyright Gray Areas and Data Anxiety
The Copyright Compliance Dilemma of AI Training Data
Obtaining training data by purchasing physical books and then scanning them is, to some extent, a strategy for AI companies to mitigate copyright risk—"I bought this book" seems to provide a certain legitimacy. But whether using book content to train commercial models constitutes fair use remains a highly uncertain legal gray area.
Section 107 of U.S. copyright law defines the "Fair Use" doctrine with four consideration factors: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used in relation to the whole, and the effect on the potential market for the work. Whether AI training constitutes fair use remains unsettled. Between 2023-2024, multiple landmark lawsuits are progressing: The New York Times v. OpenAI, Authors Guild v. OpenAI, Getty Images v. Stability AI, and others. Supporters cite the Google Books case (Authors Guild v. Google, 2015) as precedent, arguing that digitizing entire books to build a search index constitutes "transformative use." Opponents counter that AI models can generate text highly similar to training data, directly harming the market value of original works. The EU's AI Act and Digital Single Market Copyright Directive have taken a different regulatory path, requiring explicit copyright exception provisions.
Behavioral Deviation Driven by Data Anxiety
This phenomenon fundamentally reflects the anxiety and disorder pervading the entire AI industry regarding data acquisition. When data is viewed as a core competitive advantage, the means of obtaining it may breach social norms and cultural boundaries that were previously taken for granted. The destruction of rare books is just one extreme manifestation, reminding us to consider: in AI's headlong rush forward, are we paying costs that haven't been adequately discussed in exchange for marginal improvements in model capabilities?
Notably, this data anxiety is not without technical alternatives. Synthetic data generation, data augmentation techniques, and more efficient training algorithms (such as reducing total data requirements through improved data quality filtering) are all active research directions. However, in the current competitive landscape, these long-term technical solutions often give way to the short-term impulse to hoard data.
Conclusion: Data Can Be Copied, but Shredded Rare Books Cannot Be Restored
The topic "AI companies are shredding rare books" resonates so deeply because it transforms the abstract concept of "data acquisition" into a disturbing image: a guillotine blade falling on a book that may be one of a kind. The tension between technological progress and cultural preservation is starkly highlighted here.
Perhaps we cannot stop the broader trend of book digitization, but at the very least, we should establish mechanisms for identifying and protecting rare documents, subjecting irreversible acts of destruction to proper constraints. Data can be copied, but shredded rare books cannot be restored.
Key Takeaways
Related articles

MathCode: An AI Coding Agent Built Specifically for Mathematical Computation
Deep dive into MathCode, an AI coding Agent for math computation. Learn how it uses code execution to overcome LLM reasoning limitations for precise symbolic and numerical calculations.

The Dilemma and Way Forward for Formal Verification: Lessons from 50 Years of Debate
Revisiting the 1979 DeMillo critique of formal verification: examining whether modern tools like Coq, TLA+, and Lean solve fundamental issues of specification correctness and social processes.

In-Depth Analysis of the St. Lucie Nuclear Power Plant Unit 1 Manual Shutdown Event
Detailed analysis of the St. Lucie Unit 1 manual shutdown event, covering 3 control rods dropping into the core, PWR safety mechanisms, and defense in depth principles for nuclear safety.