Rare Books Traced to AI Training Facilities: The Data Copyright Dispute and the Industry's Transparency Crisis

Rare books traced to Amazon's AI facility expose the growing tension between AI training data needs and copyright.
A report tracking rare books to an Amazon AI training facility has reignited debate over AI training data sourcing and copyright. As high-quality digital text becomes scarce, tech companies are turning to physical books — especially rare ones — for unique corpus value. The article examines the legal gray area between fair use and infringement, contrasts AI training with the Google Books precedent, and highlights the industry's critical lack of training data transparency alongside emerging regulatory and licensing frameworks.
The Mysterious Journey of a Shipment of Rare Books
Recently, a report from the Hacker News community sparked widespread discussion about the origins of AI training data. The headline cut straight to the heart of the matter: We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility.
This investigative report touches on one of the most sensitive issues in the development of generative AI — the sourcing and legality of training data. When rare, out-of-print, and even historically significant physical books are shipped in bulk to a tech giant's data processing facility, people can't help but ask: What exactly happened to these books? How will their content be used?

Why Physical Books Have Become a Target for AI Training
From Digital Text to Physical Scanning
In the training of large language models, text data is the most essential "fuel." Early models primarily relied on publicly available digitized text from the internet — Wikipedia, news websites, open-source code repositories, and the like. However, as model scales continue to expand, high-quality public text resources are rapidly being exhausted.
This phenomenon is known in the industry as the "data wall" problem. According to the Chinchilla scaling laws proposed by the DeepMind team in 2022, the number of tokens needed to train an optimal model should be proportional to the model's parameter count — a model with hundreds of billions of parameters theoretically requires trillions of tokens of high-quality training data. According to estimates by the research organization Epoch AI, the total volume of publicly accessible high-quality text data on the internet falls somewhere between 4.6 trillion and 17 trillion tokens, and the most advanced current models are approaching or have already exceeded this upper bound. While synthetic data (training data generated by AI models themselves) is seen as a potential alternative, research has shown that training purely on synthetic data can lead to "model collapse" — a progressive degradation in the diversity and quality of model outputs across generations. As a result, authentic, human-created high-quality text remains irreplaceable.
It is widely acknowledged in the industry that high-quality book text — especially professionally edited publications with rigorous structure and high knowledge density — offers irreplaceable value for improving a model's linguistic capabilities and depth of knowledge. Once easily accessible digital text has been consumed, tech companies have begun turning their attention to physical books that have yet to be digitized.
The Unique Corpus Value of Rare Books
The report specifically mentions "rare books." These types of books often contain unique content that is difficult to find online: out-of-print works, in-depth materials from specialized fields, historical documents, and more. For AI training that pursues data diversity and scarcity, this content is precisely the kind of "fresh corpus" that holds enormous appeal.
Scanning physical books, performing OCR recognition, and converting them into trainable text is becoming a method some organizations use to supplement training data. Modern OCR (Optical Character Recognition) technology is quite mature — deep learning-powered OCR systems can achieve over 99% accuracy on printed text, and can produce high-fidelity results even with aged publications or unusual fonts. It's worth noting that using book corpora for AI training is not entirely new — the previously exposed Books3 dataset contained approximately 196,000 ebooks obtained through shadow libraries, and was used by multiple AI companies for training, triggering large-scale copyright lawsuits. Compared to obtaining data from gray-area electronic sources, directly purchasing and scanning physical books offers more "controllability" at an operational level, but that doesn't mean it's legally safe.
While this process is not technically complex, the copyright and ethical issues it involves are quite thorny.
The Gray Area of AI Training Data Copyright
Fair Use or Copyright Infringement?
Copyright disputes surrounding AI training data have been intensifying globally in recent years. Multiple publishers and writers' organizations have filed lawsuits against leading AI companies, with the core point of contention being: does using copyrighted works to train commercial AI models without authorization constitute infringement?
AI companies typically invoke the "fair use" doctrine in their defense, arguing that the training process constitutes "transformative use" of the original content. Section 107 of U.S. copyright law establishes four factors for determining fair use: (1) the purpose and character of the use, including whether it is commercial or transformative in nature; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the whole; and (4) the effect of the use on the potential market for the original work. AI companies emphasize that the training process "transforms" original text into statistical patterns and weight parameters, and that the model doesn't "store" the original text, making it a highly transformative use. Rights holders, however, argue that the model is essentially extracting commercial value from protected creative works and should require authorization or compensation.
Several major lawsuits are currently underway worldwide. The New York Times filed a lawsuit against OpenAI and Microsoft in late 2023, alleging unauthorized use of millions of news articles for model training. Thousands of authors (including prominent names like John Grisham and George R.R. Martin) filed a class-action lawsuit against OpenAI through the Authors Guild. The outcomes of these cases will largely determine the legal boundaries for AI training data use. It's also worth noting that legislative attitudes vary significantly across countries: Japan's copyright law, revised in 2018, explicitly permits the use of copyrighted content for machine learning purposes; the EU's Digital Single Market Copyright Directive allows text and data mining but grants rights holders the option to "opt out."
This rare books tracking report attracted attention precisely because it made an abstract copyright dispute tangible — no longer an invisible scraping of text from the internet, but a physically traceable, very real shipment of books.
Legal Risks of Physical Scanning for AI Training
Scanning physical books is nothing new in itself. The Google Books project scanned library collections on a massive scale years ago, which embroiled it in a decade-long legal battle. Ultimately, the court ruled that scanning books and providing snippet search results constituted fair use.
The reasoning behind the Google Books case (Authors Guild v. Google, 2015) offers important reference points. The Second Circuit Court of Appeals determined that Google's scanning of books and providing limited snippets was "highly transformative" because it created an entirely new function — search and discovery — rather than replacing the reading experience of the original work. The court also specifically noted that because only limited snippets were displayed, the activity would not substantially substitute for sales of the original works. However, the circumstances of AI training differ from this in critical ways: the training process typically requires ingesting the complete content of a work (even if the final model doesn't store it verbatim), and the trained model can generate output that is similar in style, structure, and even content to the original work, potentially creating direct market competition.
Therefore, using scanned content to train AI models capable of generating competitive content is fundamentally different from providing search snippets. The former may directly impact the market value of original works, making its legal classification far more complex and uncertain than the Google Books case. The fourth factor of the fair use test — "the effect on the potential market for the original work" — will likely become the most contested focal point in such cases.
The Industry's Dilemma of Missing Training Data Transparency
Training Data Composition Remains a "Black Box"
The training data composition of most mainstream AI models today is highly opaque. Driven by commercial competition and potential legal risks, companies are often reluctant to disclose their specific data sources, making external oversight nearly impossible.
This lack of transparency stands in stark contrast to the best practices advocated by the academic community. As early as 2018, researchers proposed the "Datasheets for Datasets" initiative, arguing that every dataset used for machine learning should come with detailed documentation describing its creation motivation, data composition, collection process, preprocessing methods, and potential biases. Google later introduced "Model Cards," also recommending public disclosure of model training data sources and evaluation results. In practice, however, OpenAI's GPT-4 technical report devoted only a few sentences to describing training data, citing "the competitive landscape and safety implications" as reasons for withholding details. Meta's LLaMA series, while more open, has also frequently been criticized for overly vague data source disclosures. Only a few fully open-source projects (such as EleutherAI's The Pile dataset) have provided relatively thorough documentation of training data composition.
Precisely because of this, "detective-style" investigative methods like physically tracking shipments have become one of the few avenues for outsiders to glimpse AI data sources. This in itself reflects a serious lack of industry transparency and demonstrates that public concern about the origins of AI training data continues to intensify.
The Urgent Need for Industry Rules and Compensation Mechanisms
As incidents like these continue to come to light, establishing clearer industry rules has become an urgent necessity. Possible directions include: mandatory disclosure of training data sources, licensing mechanisms for copyrighted content, and fair compensation schemes for original creators.
On the regulatory front, various parties have already begun taking action. The EU's AI Act officially came into force in 2024, requiring providers of General-Purpose AI (GPAI) systems to provide a "sufficiently detailed summary" of training data, including descriptions of the training datasets used and their sources. The U.S. Copyright Office also held multiple public hearings and public comment periods between 2023 and 2024 on AI and copyright issues, seeking to clarify how the existing legal framework applies in the age of AI. On the industry practice side, some organizations have begun exploring commercial models for data licensing: OpenAI has reached content licensing agreements with the Associated Press, Axel Springer (parent company of Die Welt), France's Le Monde, and other media organizations; Shutterstock and Getty Images have each launched image licensing programs for AI training and pledged to share revenue with contributors. Reddit licensed its platform content to Google for AI training at approximately $60 million per year.
These explorations may represent a more sustainable path forward — ensuring that content creators receive their rightful share of the value created by AI. However, challenges remain: how to provide retroactive compensation for data that has already been used, and how to ensure small and independent creators have fair standing in negotiations — these questions still await more systematic solutions.
Conclusion: An Inescapable Data Ethics Question for the AI Era
The shipping trail of a batch of rare books has unexpectedly illuminated a long-overlooked corner of the generative AI supply chain. It reminds us that no matter how advanced the technology becomes, AI's intelligence is ultimately built upon the knowledge accumulated by humanity.
When these books — vessels of humanity's intellectual heritage — are fed into training facilities, we need to seriously consider: how can we advance technology while respecting the rights of knowledge creators and maintaining a healthy knowledge production ecosystem? This is not merely a legal question but a fundamental ethical proposition for the AI era.
It should be noted that this article is based on analysis of a single-source report from the Hacker News community, and the specific details of the report and formal responses from the parties involved remain to be further verified. Nevertheless, the discussion it has sparked about the sourcing of AI training data undoubtedly carries universal real-world significance.
Related articles

New MCP Release: How Stateless Protocol Is Reshaping AI Tool-Calling Architecture
MCP's new version introduces stateless protocol design for better scalability and reliability. A free 5-hour livestream on Sept 9 covers protocol evolution, server building, and the AI agent ecosystem.

Fable 5 vs Opus 5: A Hands-On Comparison of AI-Generated 2D Sprites
Comparing Claude Fable 5 and Opus 5 generating 2D knight sprites with identical prompts — analyzing file count, animations, technical approach, and cost.

Qwen3.8 27B Scores 52 Points: How a Mid-Size Open-Source Model Is Rewriting the Performance Landscape
Alibaba's Qwen3.8 27B scores 52 on Artificial Analysis, rivaling flagship models with just 27B parameters. Explore its performance, local deployment advantages, and impact on the open-source model landscape.