Microsoft Internally Called AI Data Scraping 'The Largest Theft of Labor in Human History'

Microsoft internally labeled AI training data scraping as 'labor theft,' clashing sharply with its public fair use claims.
Newly unsealed court documents reveal that Microsoft internally described AI training data collection as "theft," with one executive calling it "the largest theft of labor in human history." The documents show both Microsoft and OpenAI scraped content from behind The New York Times' paywall, while internal teams warned the practice would "gut" publishers — a stark contradiction to their public fair use defenses. The findings could accelerate content licensing deals across the AI industry and push courts toward clearer rulings on training data legality.
Internal Documents Expose an Industry Secret
A set of newly unsealed court documents is pushing the controversy over AI training data into the spotlight. According to reports, these unredacted litigation materials reveal that Microsoft privately described OpenAI's data practices as "theft" — even as both companies were scraping content from behind The New York Times' paywall to build training datasets.
Perhaps the most striking detail in the documents is a single phrase: a Microsoft executive reportedly characterized AI data scraping as "the largest theft of labor in human history." The remark has drawn attention not only for its blunt language, but because it came from inside a tech giant deeply embedded in the AI boom — not from an outside critic.

The Gray Zone Between Paywalled Content and Training Data
The central finding in the documents is this: both Microsoft and OpenAI scraped content from behind The New York Times' paywall and incorporated it into training datasets. Paywalls are a publisher's primary mechanism for protecting the value of their content and sustaining subscription revenue. When that protected content is harvested at scale to train large language models, it strikes at the most commercially sensitive nerve in the media industry.
Notably, the documents also indicate that employees at both companies internally warned that these practices would "gut" publishers. In other words, the technical teams involved were not unaware of the potential consequences — they pressed forward with full knowledge that their actions could harm content producers. This gap between internal warnings and external justifications is precisely the kind of evidence that plaintiffs find most valuable in litigation of this nature.
From a technical standpoint, AI companies typically access paywalled content through several routes: leveraging search engine cached pages, purchasing third-party datasets that were compiled before paywalls were in place or that circumvented them, or conducting bulk scraping before paywall restrictions tightened. Common Crawl — one of the most widely used public datasets in AI training — contains vast historical archives of content that was scraped before paywalls existed or when technical loopholes were present. The GPT model series and numerous open-source LLMs rely heavily on Common Crawl as a data source. This means that even if an AI company never directly accessed a paywall, its training data may still indirectly include protected content — one reason why copyright litigation in this space is extraordinarily complex from an evidentiary standpoint.
Why These Documents Matter
AI companies have consistently taken a unified public position in copyright lawsuits: that using training data constitutes "fair use" under copyright law, and that scraping publicly accessible web content is standard industry practice. The value of these unredacted documents lies in the possibility that they reveal internal thinking that conflicts with those public defenses.
When a company argues in court that its conduct was lawful, while internal emails or memos use words like "theft" to describe that same conduct, the contradiction can significantly undermine the credibility of its legal defense. For plaintiffs like The New York Times, such internal characterizations are close to ideal evidence.
A note on fair use: "Fair use" is a core affirmative defense under U.S. copyright law that permits the unauthorized use of copyrighted works under certain conditions. Courts typically weigh four factors: the purpose and character of the use (whether it is transformative), the nature of the original work, the amount used, and the effect on the market for the original. The prevailing AI industry argument is that model training constitutes "transformative use" — extracting statistical patterns from large volumes of text rather than reproducing the content itself, analogous to the way humans learn through reading. However, when a model can reproduce specific passages from its training data verbatim (the so-called "memorization" phenomenon), or when it directly substitutes for the market demand served by original content, the transformative use defense faces serious challenges. In the New York Times v. OpenAI case, the plaintiff submitted test screenshots showing the model reproducing lengthy excerpts from paywalled articles — a targeted attack on precisely this vulnerability.
Deeper Implications for the AI Industry
This episode reflects a structural tension in generative AI that remains unresolved: large models depend heavily on vast quantities of high-quality text, much of which is the product of professional content creators' labor. The characterization of "labor theft" is fundamentally a challenge to the sustainability — legally and ethically — of a system in which AI absorbs years of output from journalists, editors, and authors without compensating them.
For the industry as a whole, these documents may trigger two cascading consequences. First, they could accelerate the adoption of content licensing agreements, as more AI companies find themselves compelled to negotiate paid deals with publishers rather than relying on scraping. Second, they may push regulators and courts toward clearer definitions of what constitutes lawful use of training data.
Where things currently stand: Some AI companies have already begun securing training data through paid licensing arrangements to mitigate legal exposure. OpenAI has signed data licensing agreements with the Associated Press and Axel Springer; Google struck a deal with Reddit to use its forum content for training. Critics point out, however, that the licensing fees involved are far below the actual commercial value of the content, and that agreements tend to favor large media organizations with negotiating leverage — leaving a vast number of smaller and independent content creators without protection. Meanwhile, some publishers have updated their robots.txt files to explicitly prohibit AI crawlers, but this technical convention carries no legal force and relies entirely on voluntary compliance from AI companies.
Closing Note
It is worth noting that this article is based on a single news source; the original court documents, the identities of the executives involved, and the full context of the statements cited have yet to be independently verified. Regardless of how the litigation ultimately resolves, however, the phrase "the largest theft of labor in human history" — uttered inside one of the world's most powerful technology companies — has already added a footnote to the debate over AI data ethics that is difficult to ignore. It is a reminder to the industry that the legitimacy of technological progress must ultimately be grounded in a balance between the rights of content creators and the expansion of AI capabilities.
Related articles

Opus 5's Ethical Boundaries: From Refusal to "Horror-Themed Project" — An Accidental Jailbreak Experiment
A developer bypassed Claude Opus 5's refusal by renaming a fruit fly simulation a "horror-themed project." Explore what this reveals about LLM content moderation and AI alignment.

Is Voice AI Actually Reliable in Real-World Call Center Scenarios?
Can Voice AI really handle real call center chaos — interruptions, noise, and intent shifts? We break down the technical limits, demo traps, and how to evaluate reliability.

The Rogue AI Agent Problem: Can AI Supervising AI Be the Cure?
As AI agents outpace human review capacity in speed, duration, and scale, enterprises face a critical oversight gap. Can AI supervising AI be the fix?