AirTag Tracking Reveals: Amazon Suspected of Destroying Rare Books to Train AI

AirTag tracking reveals Amazon may be destroying rare books to obtain AI training data.
An investigation by 404 Media used Apple AirTags to track rare books sold in bulk, revealing they ended up at an Amazon AI training facility in Las Vegas. The findings suggest AI companies may be using destructive scanning to convert irreplaceable rare books into training data, raising serious concerns about copyright infringement, cultural heritage destruction, and the ethical boundaries of AI data acquisition in an era of increasing data scarcity.
An Investigation Launched with an AirTag
When tech journalists slipped an Apple AirTag into a rare book, they probably didn't expect that this tiny tracker would expose a little-known corner of AI training. The Apple AirTag is a small tracking device based on Ultra-Wideband (UWB) technology and Bluetooth Low Energy protocol. It leverages Apple's "Find My" network—a crowdsourced location system composed of over one billion Apple devices worldwide—to achieve continuous location tracking of items. Any nearby Apple device anonymously relays the AirTag's encrypted location signal to the cloud in the background, enabling precise global positioning even though the tracker itself has no GPS or cellular network connection.
According to an investigation by tech news outlet 404 Media, a journalist arranged with a book dealer to place an AirTag inside a rare book that would be shipped as part of a bulk order. 404 Media is an independent tech news organization founded in 2023 by four veteran journalists formerly of Vice's Motherboard tech channel, known for in-depth investigative reporting focused on technology's impact on society. Its reader-subscription business model, free from advertising revenue, provides a degree of editorial independence. The tracking data showed that the book ultimately arrived at an Amazon AI training facility in Las Vegas, Nevada.
This discovery has attracted widespread attention because it touches on one of the most sensitive topics in today's AI industry: where does training data come from, and at what cost is it obtained? If the investigation holds true, then a batch of collectible physical books may be getting dismantled, scanned, and then destroyed—merely to satisfy the enormous appetite of large language models for text data.

Why AI Companies Need Book Data
High-Quality Text Resources Are Increasingly Scarce
To understand the context of this event, one must first understand the unique value of books for AI training. The capabilities of large language models depend heavily on the quality of their training data. Current mainstream large language models have training data scales reaching the trillions of tokens—GPT-4, for example, is estimated by industry observers to have been trained on over 13 trillion tokens. However, the total amount of high-quality English text available on the internet is finite. Research institution Epoch AI warned in a 2022 study that high-quality language data could be "exhausted" by around 2026. This phenomenon, known as the "data wall," is forcing AI companies to seek new data sources, including synthetic data generation, multimodal data conversion, and digitizing text resources from the physical world.
Compared to the fragmented, low-quality content flooding the internet, books represent high-quality text from human knowledge that has been edited, structurally complete, and logically rigorous. Whether literary works, academic monographs, or professional reference books, the extended coherent narratives and in-depth arguments found in books are precisely the materials models need to learn complex reasoning and language expression. Research shows that models trained on high-quality book corpora significantly outperform those trained solely on web data in logical reasoning, long-form text generation, and factual accuracy. This is why numerous AI companies are finding ways to obtain book data—from public book datasets (such as Books3, which contains approximately 196,000 books), to negotiating licensing with publishers, to what now appears to be physical book scanning and destruction.
From Physical to Digital: The Cost of Destructive Scanning
The most efficient way to convert physical books into training data is often "destructive scanning." The typical process involves: first using industrial paper cutters to remove the spine, completely separating the pages into loose sheets; then feeding the loose pages into high-speed document scanners—these industrial-grade devices can scan 100-200 pages per minute, far exceeding the efficiency of manual page-turning scanning; the scanned images are then converted into editable text data through Optical Character Recognition (OCR) technology. By comparison, non-destructive scanning requires operators to turn and align each page individually, typically processing only a few pages per minute—an efficiency gap of several dozen times. The Google Books project also extensively used destructive scanning methods in its early stages, but it typically processed ordinary holdings from libraries with multiple copies, not unique or rare editions.
For ordinary books, this may be unremarkable, but when the subjects being processed are rare or even unique volumes, the nature of the issue changes entirely—it means irreplaceable cultural carriers are being completely destroyed in the digitization process.
What Controversies Has This Practice Sparked
The Legal Gray Zone of Copyright and Data Sourcing
This incident has once again pushed the legality of AI training data to the forefront. Legal disputes surrounding AI training data are already unfolding across multiple jurisdictions worldwide. In the United States, the core controversy centers on the scope of the "Fair Use" principle—Section 107 of U.S. copyright law permits unauthorized use of copyrighted works in contexts such as commentary, education, and research, but whether commercial AI training constitutes "transformative use" remains highly contested. In recent years, from The New York Times suing OpenAI (alleging ChatGPT can reproduce its paid articles nearly verbatim) to class-action lawsuits filed by numerous authors and artists, the debate over "whether AI companies have the right to use copyrighted works for training" has never ceased. In the EU, the AI Act and the Digital Single Market Copyright Directive provide rights holders with an "opt-out" mechanism; Japan has taken a relatively permissive stance, with its 2018 copyright law revision explicitly allowing the use of copyrighted materials for machine learning purposes.
Even if a book is legally purchased, whether using its contents for commercial AI model training constitutes infringement remains a legally unresolved gray area. Under the "First Sale Doctrine" in U.S. copyright law, a purchaser has the right to resell or dispose of their purchased physical copy, but this right does not extend to copying the content and commercially reusing it. Owning a book does not equate to obtaining authorization to use its contents for any commercial purpose, especially when such use may directly conflict with the original author's commercial interests.
Cultural Heritage Faces Irreversible Loss
If copyright issues still have room for debate, the cultural loss from destroying rare books is even more heartbreaking. In the fields of antiquarian book collecting and library science, the value assessment of "rare books" is based on multiple dimensions: edition scarcity (number of extant copies), historical significance (first editions, author-signed copies, important annotated copies), physical condition (evidence of binding craftsmanship, paper materials, and printing technology of the era), and provenance (ownership transfer history). A first-edition book from the 19th century may be worth far more than its textual content alone—it serves as physical evidence of a specific era's printing technology, design aesthetics, and cultural dissemination methods.
Rare books are precious not only for their textual content but also for their historical value, edition value, and collectible value as physical objects. The Antiquarian Booksellers' Association of America (ABAA) has strict authentication standards for rare books, and many rare books are irreplaceable cultural heritage. Once cut apart and destroyed, these values are permanently lost with no possibility of recovery.
Trading an irreplaceable cultural artifact for a piece of digital data that can be infinitely copied strikes many as putting the cart before the horse. It reflects a concerning tendency in the current AI race: in pursuit of data scale and model performance, some participants may be ignoring the longer-term cultural and ethical costs.
Deeper Implications of This Incident
Where Are the Ethical Boundaries in the Age of Data Hunger
The significance of this investigation extends beyond the actions of a single company. It reveals a reality that is rapidly approaching: as high-quality training data becomes increasingly scarce, AI companies' methods of obtaining data may become increasingly aggressive. When public data has been "squeezed dry," physical-world resources may become the next target. The industry has already observed multiple strategies for dealing with the "data wall": some companies invest in synthetic data generation technology, using AI itself to produce training data; others sign expensive licensing agreements with content platforms like Reddit and news organizations; still others explore converting video, audio, and other multimodal data into text. Directly digitizing physical books may be the crudest and most controversial approach among them.
The fact that journalists used a consumer-grade tracking device like the AirTag to expose practices behind the AI supply chain is itself highly symbolic—technological tools in ordinary people's hands are becoming instruments for holding tech giants accountable. This investigative approach recalls classic investigation cases where environmental organizations track illegal fishing fleets or journalists trace e-waste export routes. It also reminds us that the AI industry's transparency issues urgently need more sustained external oversight.
Industry Standards and Regulatory Frameworks Urgently Need Establishment
This incident calls for clearer industry standards and regulatory frameworks. Consensus should be built on at least the following levels:
- Training data sources should be traceable and disclosable
- Culturally valuable physical materials should prioritize non-destructive methods during digitization
- Special protection mechanisms should be established for rare and out-of-print books
Currently, some countries and regions have begun taking steps in this direction. The EU's AI Act requires high-risk AI systems to disclose training data summaries; the U.S. Congress is also discussing multiple bills related to AI training data transparency. But how to balance cultural heritage protection with AI training data acquisition remains an area where no mature regulatory framework has been established globally.
Interestingly, this investigation currently comes primarily from a single media outlet, 404 Media, and the companies involved have not yet made a comprehensive response. The specific scale and whether this constitutes systematic behavior still require further verification. But whether this is an isolated case or a widespread phenomenon, the discussions it has sparked deserve serious attention from the entire industry.
Conclusion
From a tiny AirTag, to an AI training facility, to a destroyed rare book—this tracking chain connects the ethical dilemmas of data acquisition in the AI era. Technological progress is certainly important, but when it comes at the cost of irreversible cultural loss, we need to pause and ask: is this cost truly worth it? On the path to pursuing more powerful AI, how to uphold the bottom lines of cultural heritage protection and copyright ethics will be an unavoidable challenge for all practitioners.
Related articles

4 Core Skills More Valuable Than Writing Code in the AI Era
When AI can efficiently write code, where does a developer's competitive edge lie? This article breaks down 4 skills more valuable than coding in the AI era.

Don't Buy the "Tech Is Dead" Lies: Java, Web Dev, and DSA Are Alive and Well
Debunking claims like "Spring Boot is dead" and "Web dev is dead." Job market data proves these technologies thrive. Learn how AI reshapes—not replaces—developers.

AI Agent Learning Roadmap: A Four-Stage Guide from Zero to Production
A complete AI Agent learning roadmap covering four stages—foundations, core frameworks, hands-on projects, and advanced mastery—to help beginners build production-ready agents in six months.