AI Companies Destroying Rare Books for Training Data: The Cultural Cost Behind Efficiency

AI companies are destroying rare books for training data, raising cultural preservation concerns.
Some AI companies are purchasing rare antique books, destructively scanning them for model training data, and then discarding the originals—even near-unique copies. While this approach maximizes digitization efficiency, it permanently destroys irreplaceable cultural artifacts whose value extends far beyond their text. The practice highlights a growing tension between AI's insatiable data demands and cultural heritage protection, compounded by legal gray areas and a lack of regulatory oversight.
An Overlooked Part of AI Training
When we talk about training data for large language models, the focus tends to be on web text, code repositories, and public datasets. Yet a little-known and deeply unsettling reality is emerging: some AI companies are purchasing large quantities of antique books and rare printed works, scanning their contents for model training, and then physically destroying these books—even when only a handful of copies remain in the entire world.
This phenomenon has recently sparked widespread discussion on communities like Reddit. It touches on the deep tension between technological progress and cultural heritage preservation, and compels us to reexamine the ethical costs hidden behind the seemingly neutral technical process of "data acquisition."

Why AI Companies Choose to Destroy Books for Training Data
The Industrial Logic of Destructive Scanning
To understand this practice, you need to grasp the technical realities of large-scale book digitization. Traditional non-destructive scanning requires manual page-turning or expensive non-destructive equipment—it's slow and costly. Non-destructive scanning technologies include V-shaped book cradle scanners (such as the Zeutschel OS series), overhead aerial photography systems, and multispectral imaging devices designed specifically for precious manuscripts. These can capture high-quality digital images without touching or damaging the originals, but processing a single volume typically takes hours to days, and the equipment is prohibitively expensive.
Destructive scanning is an entirely different matter: the spine is cut, the book is disassembled into individual pages, and then processed in bulk through high-speed automatic document feeders (such as Fujitsu fi-series or Kodak Alaris industrial-grade devices), handling tens to hundreds of loose pages per minute—an efficiency gain of several orders of magnitude.
For AI companies pursuing massive training corpora, speed and scale are the top priorities. Current large language models demand astronomical amounts of text data—GPT-3 was trained on approximately 300 billion tokens, and GPT-4 and subsequent models have scaled up by multiples. In a 2022 study, research organization Epoch AI warned that the available high-quality text data on the internet could be "exhausted" around 2026—meaning conventional sources like web text, academic papers, and open-source code would no longer be sufficient to support training even larger models. This "data wall" effect has driven AI companies to turn their attention to previously undigitized physical books, private archives, and historical documents.
When the goal is to convert millions of books into trainable text data as quickly as possible, destructive scanning becomes the "rational" choice from a pure engineering efficiency standpoint. This stands in stark contrast to the Google Books project—when Google launched it in 2004, the company partnered with institutions like the University of Michigan and Stanford University and insisted on using custom-built non-destructive scanning equipment. Progress was slower, but the physical integrity of the books was preserved.
Copyright Evasion and the Gray Area of "Physical Elimination"
Even more alarming is that destroying the originals may also serve as a strategy for mitigating legal risk. Under U.S. copyright law, the "First Sale Doctrine" (17 U.S.C. § 109) allows the lawful owner of a copyrighted work's copy to resell, lend, or destroy that copy, but it does not automatically grant the right to reproduce or create derivative works. When AI companies digitize books for model training, this fundamentally involves an act of "reproduction"—a distinct legal issue from mere physical disposal.
Whether AI training constitutes "Fair Use" under copyright law remains a fiercely debated legal question—lawsuits such as The New York Times v. OpenAI and Getty Images v. Stability AI are proceeding simultaneously across multiple jurisdictions. The earlier Books3 dataset incident (which contained approximately 196,000 pirated ebooks) already revealed the industry's aggressive tendencies in data acquisition, and purchasing and destroying physical antiquarian books represents a further escalation of that tendency.
By physically eliminating the originals, companies may be attempting to obscure the traceability of data sources—if the original book no longer exists, it becomes much harder to trace exactly which works were used for training, and much harder to prove the specific scope and extent of infringement, thereby creating defensive room for potential copyright disputes.
What Irreversible Cultural Losses Does Destroying Rare Books Entail?
Rare Antique Books Are Disappearing
The most controversial aspect is this: among the books being destroyed are many antique volumes that exist in extremely limited numbers—some nearly unique copies. A physical book that carries the printing techniques, binding styles, and marginal annotations of a specific historical period holds value far beyond its textual content alone.
In the fields of cultural preservation and book history research, a physical book carries information far exceeding its written words. The fiber composition and manufacturing process of the paper can reveal its production era and origin—for example, the use of cotton-fiber paper versus wood-pulp paper marks different eras of papermaking. Binding methods reflect the publishing traditions of specific regions and periods. Watermarks are key evidence for identifying paper provenance. Marginalia and bookplates record a book's ownership history (provenance), constituting invaluable scholarly resources. Fermat's famous "Last Theorem" was famously preserved as a marginal note on a book page. Additionally, flowers, letters, newspaper clippings, and other "ephemera" tucked between pages constitute unique historical records in their own right.
Digitization can only preserve textual information—it cannot preserve the texture of the paper, the composition of the ink, handwritten notes between pages, bookplates, or the full range of historical information a book carries as a material cultural artifact. While multispectral imaging can reveal certain ink traces invisible to the naked eye, it still cannot fully reproduce the three-dimensional physical characteristics and material properties of the original object. When the last remaining copies are disassembled and destroyed, information across all these dimensions is permanently lost—irrecoverable by any scanning technology.
Treating Rare Books as Tokens: Data Is Not Knowledge
This phenomenon reveals a deeper cognitive bias: equating "data" with "knowledge" or even "culture." Through the industrialized lens of AI training, a precious antiquarian book is reduced to a pile of extractable tokens.
From a technical standpoint, this reduction has a clear mechanism: in the training pipeline of large language models, raw text first passes through a tokenizer, which segments it into a series of tokens—these may be complete words, word roots, character fragments, or even punctuation marks. Taking the BPE (Byte Pair Encoding) tokenization method used by the GPT series as an example, a passage of text is converted into a purely numerical sequence, and all visual information—typographic formatting, font choices, page layout—is completely discarded in the process. For rare books, this means handwriting characteristics of manuscripts, the line quality of woodblock illustrations, the uneven ink distribution of movable type printing—all of this information is obliterated during tokenization.
This dimensional reduction—from a rich, multidimensional material artifact to a one-dimensional numerical sequence—is the technical reality behind the criticism of "treating rare books as tokens." This reductionist approach to value assessment ignores the irreplaceable independent value of physical artifacts in historical research, authentication, art appreciation, and beyond.
Where Is the Line Between Technical Efficiency and Ethics in AI Training?
Efficiency Should Not Be the Only Measure
It must be acknowledged that destructive scanning is not unique to the AI era. Libraries and archival institutions also employ this method when processing large batches of low-value printed materials with many existing copies. The crux of the controversy lies not in the technology itself, but in the choice of what it's applied to—when destructive scanning is indiscriminately applied to rare antiquarian books, the efficiency-first logic crosses a reasonable ethical boundary.
AI Data Collection Lacks Transparency and Oversight
Currently, the public knows very little about the specific data collection processes of AI companies. Which books are being purchased? What scanning methods are used? What happens to the originals? Most of this information remains opaque. In the absence of industry standards and external oversight, commercial incentives around cost and efficiency easily override considerations of cultural preservation.
From an institutional perspective, the world's major cultural heritage protection legal frameworks—such as UNESCO's Convention Concerning the Protection of the World Cultural and Natural Heritage (1972) and various national cultural property laws—primarily target buildings, archaeological sites, and museum collections. Protection for privately held antique books and rare printed works is comparatively limited. In most jurisdictions, it is entirely legal for a private individual to purchase and destroy a book they own, even if it is the last surviving copy. The EU AI Act (effective 2024), while imposing tiered regulatory requirements on AI system development and use, focuses on algorithmic transparency and high-risk application scenarios—it does not specifically address cultural heritage protection during the training data collection process. This regulatory vacuum means that, driven by commercial interests, the destructive exploitation of rare cultural resources faces no effective institutional constraints.
This also raises an urgent question: as we push the boundaries of AI capabilities, have we established sufficient mechanisms to ensure that technological development does not come at the cost of irreversible cultural loss?
Reexamining the "Raw Material Cost" of AI
This issue reminds us that AI's "intelligence" does not emerge from nothing—behind it lies an enormous data acquisition chain, and that chain is producing real-world consequences that have not been adequately discussed. When training data sources involve irreplaceable cultural heritage, what we need is not just more powerful models, but serious reflection on the ethics of data acquisition.
A destroyed unique copy cannot be recovered by any future technology. Amid the race toward ever-greater AI capabilities, how to balance technical efficiency with cultural preservation is a question the entire industry cannot afford to ignore.
Key Takeaways
Related articles

Risklytics: An Insurance Brokerage Platform Built for Frontier Tech Companies in AI, Nuclear Fusion, and Beyond
YC S26 startup Risklytics provides specialized insurance brokerage for AI, nuclear fusion, and autonomous driving companies, solving the gap where traditional insurance fails to cover emerging tech risks.

Coze 3.0 Workflow in Practice: Build an Automated AI Agent in Three Steps
Learn to build AI Agents on Coze 3.0 in three steps: prompt engineering & API calls, RAG knowledge base construction, and multi-agent autonomous decision-making for low-code AI app development.

Gemini 3.5 Transcribe Explained: From Dictation to Intelligent Speech-to-Text
An in-depth look at Google Gemini 3.5 Transcribe's intelligent speech-to-text capabilities, covering contextual correction, terminology recognition, and real-world applications.