AI Companies Buy Books, Scan Them, Then Destroy Them: Pulping of Rare Books Sparks Cultural Preservation Debate

AI companies are buying, scanning, and pulping rare books for training data, sparking cultural preservation outrage.
Reports reveal that AI companies are industrially purchasing paper books—including rare antiquarian volumes—cutting off their spines for high-speed destructive scanning, extracting text for LLM training, and then pulping the remains. While legally defensible under first-sale doctrine, the practice raises serious concerns about irreversible cultural heritage loss, information privatization, and the urgent need for industry boundaries between data acquisition and preservation.
Overview: Physical Books Are Being Destroyed at Scale for AI Training
A discussion originating from Reddit is drawing widespread attention in the tech community: to train large AI models, some AI companies are systematically purchasing, scanning, and physically destroying paper books — including rare antiquarian volumes and out-of-print editions.
According to the original reports, this process has become highly industrialized: companies use hydraulic cutting machines to slice through book bindings, feed the loose pages into industrial-grade scanning equipment for mass digitization, and then pipe the extracted text content into AI training systems. Once the process is complete, the dismantled physical books are typically treated as waste paper and sent directly to paper recycling and pulping facilities.

Notably, this isn't an isolated practice by a single company — it's becoming an industry-wide norm. Used book dealers have even begun treating this AI boom as a business opportunity, proactively buying up secondhand books to resell to AI teams hungry for training data, forming a gray-market supply chain built around "book digitization and destruction."
Why Do AI Companies Choose to "Destroy" Rather Than "Preserve" Books?
From an engineering perspective, the core motivation for destroying books is scanning efficiency. Traditional non-destructive scanning requires manual page-turning, flattening each page against the spine — it's slow, expensive, and produces poor results for thick or tightly-bound volumes. "Destructive scanning," which involves removing the spine and separating the pages, allows the use of automatic document feeders for high-speed processing, improving efficiency by orders of magnitude.
It's worth elaborating on the technical details of destructive scanning. Destructive scanning is a method in document digitization that sacrifices the integrity of the original in exchange for extremely high processing speed. The core equipment is a high-speed Automatic Document Feeder (ADF) scanner. After the spine is removed using a hydraulic paper cutter or rotary trimming device, the loose pages can be pulled in one by one — like copy paper — and scanned on both sides simultaneously. Industrial-grade devices such as the Fujitsu fi series or Kodak i5000 series can process over 200 pages per minute, combined with OCR (Optical Character Recognition) software that converts images to searchable text in real time. By comparison, non-destructive scanning requires V-shaped book cradles or planetary scanners, with operators manually turning pages and adjusting focus — each page can take 30 seconds or more. This efficiency gap — from 120 pages per hour to 12,000 pages per hour — is the direct reason AI companies choose the destructive approach.
For AI companies pursuing massive volumes of high-quality corpus data, paper books represent an extremely valuable category of training data: they've undergone professional editing and fact-checking, feature standardized language and high information density, and are largely free from the problems of repeatedly crawled, inconsistent-quality web data. Compared to noise-filled web text, book corpus data provides clear benefits for improving a model's knowledge depth and expression quality.
To understand why books are so important for AI training, one needs to understand the quality stratification of training corpora. In the training pipeline of Large Language Models (LLMs), data quality directly determines the model's final performance. The industry typically classifies training corpora into several quality tiers: at the bottom is uncleaned web-crawled data (such as Common Crawl), which is enormous in volume (petabyte-scale) but riddled with ad copy, SEO spam, and duplicate content; the middle tier includes texts that have undergone some editorial process, such as Wikipedia and news websites; the top tier consists of professionally edited and fact-checked publications — namely books and academic papers. Research has shown that, given the same data volume, models trained on high-quality book corpora significantly outperform those trained purely on web data in terms of factual accuracy, logical coherence, and linguistic quality. OpenAI's GPT series, Google's PaLM, and other models have all heavily used book datasets like Books Corpus in their training data. As the Chinchilla scaling law revealed the optimal ratio between data volume and model parameters, the industry's demand for high-quality data has become even more urgent.
The Legal "Compliance" Space
Reports specifically note that this practice has room to be considered compliant under the U.S. legal framework, primarily relying on two principles:
- First-sale doctrine: Once a physical copy of a book has been legally purchased, the buyer has the right to dispose of that copy — including reselling or destroying it — and the copyright holder cannot interfere with what happens to the "physical object."
- Fair use: Under certain conditions, transformative use of works for purposes such as analysis and research may be deemed fair use.
The first-sale doctrine has deep legal roots. It originated from the U.S. Supreme Court's 1908 decision in Bobbs-Merrill Co. v. Straus and was later codified in Section 109 of the U.S. Copyright Act. The core meaning of this principle is that a copyright holder's control over a specific copy terminates after that copy's first lawful sale. The purchaser is free to resell, lend, rent (with exceptions for music recordings and software), display, or even destroy the copy without the copyright holder's consent. This principle is the legal foundation that allows secondhand bookstores, library lending systems, and flea markets to exist. However, the doctrine only covers the right to dispose of the "physical copy" and does not extend to copying and reusing the copyrighted content within it.
The determination of fair use is considerably more complex, requiring a rigorous four-factor test. Section 107 of the U.S. Copyright Act requires consideration of: (1) the purpose and character of the use, particularly whether it is "transformative" — that is, whether it creates new meaning or value rather than simply substituting for the original work; (2) the nature of the copyrighted work, with factual works receiving weaker protection than creative works; (3) the amount and substantiality of the portion used, where full-text copying generally weighs against a fair use claim; (4) the effect on the potential market for the original work. The key ongoing lawsuits surrounding AI training — such as Authors Guild v. OpenAI, Sarah Silverman v. Meta, and The New York Times v. OpenAI — are all fiercely debating how these four factors apply to AI scenarios. AI companies argue that the training process constitutes highly transformative use because models generate statistical patterns rather than copying original text; copyright holders counter that entire books are completely "ingested" into training systems and that model outputs directly compete with original works in the marketplace, constituting substantial infringement.
In other words, when AI companies buy books, scan them, and then destroy them, they're difficult to challenge on the level of "disposing of a physical copy they own." But this is only legality at the physical copy level — whether using scanned content for model training constitutes infringement remains the focus of multiple ongoing AI copyright lawsuits and is far from settled.
The Core Controversy: What Is the Cultural Cost of AI Progress?
What has truly ignited fierce community debate is not the legality question, but the cultural and ethical cost.
When the books being destroyed are merely copies of widely available bestsellers, the controversy is relatively limited — after all, countless identical copies exist. But when the scope extends to rare books, out-of-print editions, and antiquarian volumes with only a handful of surviving copies, the situation becomes entirely different. Each copy of such books may be an irreplaceable cultural artifact, and once cut apart and pulped, its value as a historical relic, a work of binding artistry, and a carrier of marginalia and annotations is permanently lost.
In library science and document preservation studies, a book's value extends far beyond the text content it carries. Materiality is a key concept for understanding the cultural heritage value of books. The material, thickness, and chemical composition of the paper can help date the work and trace its origin; traces of letterpress or hand-set type record the evolution of printing technology; bookplates, seal impressions, and annotations constitute a complete "provenance chain" reflecting the book's transmission path across different eras and readers. For example, handwritten notes in the margins of an 18th-century volume might document a scholar's thought process, with historical research value potentially exceeding that of the main text itself. Furthermore, a book's binding — leather covers, gilt tooling, hand-sewn stitching — constitutes physical evidence of decorative arts and craftsmanship. All these dimensions are completely erased in any form of text-extraction scanning; even high-resolution image scanning cannot fully preserve tactile qualities, scent, and three-dimensional structural information.
Digitization Does Not Equal "Preservation": The Overlooked Information Loss
A common defense is: the content has been scanned and preserved as digital text, so the disappearance of the physical object doesn't mean knowledge is lost. But this view overlooks several critical facts:
- Information loss: Scanning typically extracts only text content. The book's physical characteristics — paper, printing techniques, handwritten annotations, ownership stamps, binding — these dimensions carrying historical information are completely discarded.
- Privatization risk: Digital copies of destroyed books are held by AI companies and used for commercial model training. The public may never have access, which is fundamentally different from public preservation through libraries and archives.
- Irreversibility: Unlike high-precision conservation scanning of endangered documents, the goal of destructive scanning is efficiency rather than preservation — the result is the complete disappearance of the original.
A Deeper Industry Concern: The Conflict Between Data Hunger and Cultural Heritage
This phenomenon reflects a sharp contradiction in the current AI arms race: the insatiable demand for training data is coming into direct conflict with cultural heritage preservation.
As high-quality web text is gradually exhausted by scraping, models grow increasingly desperate for "uncontaminated" premium corpus data, and books have become one of the most attractive targets. The "data wall" currently facing the AI industry is more severe than public perception suggests. According to estimates by the Epoch AI research institute in 2022, the total volume of high-quality text data available for training on the internet is in the range of several trillion tokens, and the latest generation of large models (such as GPT-4, Llama 3, etc.) have already consumed a substantial proportion of it. At the current growth rate of model scale (where Scaling Laws require data volume to expand in tandem with parameter count), high-quality text data may be "exhausted" around 2026. This is forcing AI companies to continuously explore new data sources: synthetic data generation, multimodal-to-text conversion, and — as discussed in this article — converting not-yet-digitized paper books into training corpora. It's estimated that there are approximately 130 million unique published book titles worldwide (according to the Google Books project), a large proportion of which have never been digitized — especially out-of-print books and publications in minority languages. This enormous "untapped" corpus represents an almost irresistible temptation for data-hungry AI companies.
In the absence of clear industry standards and regulation, efficiency and cost naturally override careful consideration of cultural value.
There's a warning signal worth noting here: when destroying rare books becomes a monetizable business, market mechanisms will accelerate the process. The profit-seeking behavior of used book dealers, the data hunger of AI companies, and legal gray areas — when these three factors combine, they may cause irreversible cultural loss, the kind that often isn't truly recognized until years later.
Conclusion: Clear Boundaries Are Needed Between AI Training and Cultural Preservation
We must acknowledge that destructive scanning has its logic from technical and business perspectives, and for common books it's not necessarily unacceptable. The real problem lies in the lack of boundaries — should protective red lines be established for rare and out-of-print books? Should non-destructive scanning and public archiving be prioritized before destructive digitization? Does the industry need to establish self-regulatory standards?
This debate over antiquarian books and AI training is fundamentally asking: how great an irreversible cultural cost are we willing to pay for AI's "progress"? In an era of technology's headlong rush forward, this question deserves serious consideration from every practitioner and user.
Note: This article is based on reports circulating in the Reddit community. The specific scale of these practices and the companies involved still await verification from more authoritative sources; readers should exercise appropriate skepticism.
Related articles

GLM-5.3 Benchmark Analysis: The Globalization Journey of Chinese Large Language Models
In-depth analysis of Zhipu AI's GLM-5.3 benchmarks on Artificial Analysis, exploring third-party evaluation platforms, the GLM series evolution, and Chinese LLMs' path to global recognition.

Claude Autonomously Designs Proteins with 35% Success Rate, Far Exceeding Human Expert Performance
Anthropic's Claude achieves 35% wet-lab success rate in autonomous protein design, far surpassing the 10-15% human expert average, signaling AI's move toward real scientific productivity.

Perplexity Discover's Multilingual Support Suddenly Disappears — Why Are International Users Upset?
Perplexity Discover's multilingual news feature suddenly dropped non-English support, frustrating international users. We analyze possible causes and the broader challenges of AI product internationalization.