You Own Claude's Output, So Why Can't You Train a Model With It? The Difference Between Copyright and Contract

You own AI output by copyright, but contracts restrict you from using it to train competing models.
Anthropic claims users own Claude's output, yet prohibits using it for training rival models. This article unpacks the crucial legal distinction between copyright ownership and contractual usage restrictions, explaining why anti-distillation clauses exist across the AI industry, the massive training costs they protect, the enforcement challenges they face, and how global legal fragmentation around AI-generated content complicates the picture.
A Seemingly Contradictory Clause
On Hacker News, a developer raised a question that resonated widely: if Anthropic states in its terms of service that users "own" the output generated by Claude, then why are users prohibited from using that output to train competing models? This appears to be a self-contradictory rule, but behind it lies a deep web of copyright law, contract law, and the competitive dynamics of the AI industry.
This question doesn't just concern Anthropic. Nearly every major LLM provider—OpenAI, Google, Anthropic—includes similar anti-distillation clauses in their terms. Understanding this apparent contradiction helps us see the real landscape of the battle between AI model usage rights and intellectual property.

"Ownership" and "Usage Restrictions" Are Two Different Things
Many people conflate "ownership" with "license"—two distinct legal concepts. When Anthropic says you "own" the output, it typically means:
- You can use Claude-generated text for commercial purposes
- You can publish, modify, and distribute the content
- The company will not claim copyright over the output
However, "owning the copyright to content" does not mean "being free to do anything with it." The restrictions are imposed independently through the terms of service (a contract).
Copyright Ownership ≠ Contractual Freedom
Here's an analogy: when you buy a physical book, you own the physical object, but you can't scan the entire book and mass-produce pirated copies for sale—that's constrained by copyright law. The situation with AI output is even more nuanced: even if a company "transfers" the copyright of the output to you, it can still attach a contractual obligation—"you may not use this to train competing models"—through the terms of service you clicked to accept.
In other words, what you'd be violating is not copyright law, but the contract between you and the company. These are two entirely independent legal systems: one is intellectual property law, the other is contract law. They differ fundamentally in what they protect, how they're enforced, and what remedies they offer. Intellectual property law grants statutory rights directly from the state to creators; contract law is based on mutual consent between parties, deriving its binding force from voluntary promises. The software industry has long had precedent for this dual-track system: open-source licenses (such as GPL and MIT License) impose restrictions on how software can be used through contractual terms, even when users have access to the source code. Similarly, even if AI output isn't protected by copyright, terms of service as a contract can still restrict user behavior. The consequences of breaching a contract are typically civil damages and service termination, rather than criminal penalties—a crucial point for assessing the practical deterrent power of these clauses.
Why AI Companies Care So Much About "Model Distillation"
The core motivation behind prohibiting users from training models with output is to prevent "model distillation." This is a technique where someone queries a powerful model extensively, collects its outputs, and then uses that high-quality data to train a smaller, cheaper model.
The concept of model distillation was formally introduced by Geoffrey Hinton and colleagues in their 2015 paper Distilling the Knowledge in a Neural Network. The core idea is to transfer knowledge from a large "teacher model" to a small "student model." Specifically, the student model learns not only from standard labels but also from the teacher model's "soft labels"—probability distributions that contain rich information about relationships between categories. In the era of large language models, the meaning of distillation has expanded dramatically: attackers can query models like GPT-4 or Claude at scale via API, using the question-answer pairs as training data to produce a model with far fewer parameters but near-comparable performance. The cost of this approach can be as low as one-thousandth of the original training cost or even less—which is precisely why it poses an existential threat to model providers.
The Business Moat Built on Massive Training Costs
Training a frontier LLM requires hundreds of millions of dollars in compute, data cleaning, and human labor. Specifically, these costs break down into several components: First, compute costs—training a GPT-4-class model is estimated to require tens of thousands of NVIDIA A100/H100 GPUs running for months, with electricity and hardware rental costs alone reaching tens of millions to over a hundred million dollars. Second, data acquisition and cleaning costs—including purchasing high-quality datasets and hiring annotation teams for RLHF (Reinforcement Learning from Human Feedback), with annotator costs ranging from $15-50 per hour. While Anthropic's Constitutional AI approach reduces dependence on human annotation, it still demands substantial computational resources. Third, R&D talent costs—top AI researchers command annual salaries of several million dollars. The estimated total cost of training a frontier model has surged from approximately $10 million in 2020 to $100-500 million in 2024, with next-generation models potentially exceeding $1 billion.
If competitors could "apprentice" a model with comparable performance through API calls at minimal cost, the original model provider's business moat would collapse overnight.
This is not an idle fear. It's widely believed in the industry that the capability gains of some open-source or low-cost models were partly achieved by learning from the outputs of powerful models like GPT-4. Models like DeepSeek have also faced similar allegations. DeepSeek is a series of large language models developed by the Chinese AI company DeepSeek. In early 2024, DeepSeek-V2 drew industry attention with its extremely low inference costs and GPT-4-level performance. Researchers subsequently found through comparative testing that certain open-source models exhibited response patterns on specific tasks that closely mirrored GPT-4's, sometimes even producing the telltale phrase "As an AI language model developed by OpenAI"—considered indirect evidence of distillation. In January 2025, OpenAI publicly accused DeepSeek of potentially using its model outputs for training. This incident highlights the trend of anti-distillation clauses moving from paper constraints to real-world confrontation, while also exposing the enormous difficulties of cross-border enforcement.
Thus, anti-distillation clauses are fundamentally defensive measures by providers to protect their core assets.
The Enforcement Challenge of Anti-Distillation Clauses
Interestingly, these clauses face enormous challenges in actual enforcement. It's extremely difficult for providers to prove that a new model was "trained using our output," especially when training data has been mixed and cleaned. As such, these clauses function more as legal deterrents and bases for after-the-fact accountability, rather than absolute technical barriers.
How Users Should Understand Their Rights Boundaries
For everyday developers and enterprise users, understanding the practical boundaries of this restriction is important:
- Normal use is perfectly fine: Using Claude to write code, generate copy, or build product features—all of this falls within the authorized scope.
- The red line is "training competing models": As long as you're not using outputs in bulk to train a general-purpose LLM that directly competes with Anthropic, you won't cross the line.
- The gray area between fine-tuning and distillation: Using a small amount of output to optimize your own vertical application is fundamentally different from systematically distilling a general-purpose model, but the boundary isn't always clear.
Industry Consensus and Potential Legal Disputes
It's worth noting that the legality and enforceability of these clauses remains debated in legal circles. Some argue that if a company has fully transferred ownership of the output, then restricting its use is logically untenable. But in mainstream practice, contractual terms are still considered valid constraints.
This also reflects a fundamental unresolved contradiction across the AI industry: the copyright status of AI-generated content itself remains a gray area in most jurisdictions. In its 2023 guidance, the U.S. Copyright Office explicitly stated that content generated purely by AI without human creative input does not meet the requirements for copyright protection, because copyright law requires "human authorship." This position was upheld by the court in Thaler v. Perlmutter. In practice, however, most AI-assisted creations involve human prompt design, curation, and editing, and the copyright status of such hybrid works remains ambiguous. Even more notably, attitudes vary dramatically across jurisdictions—the Beijing Internet Court in China ruled in a 2023 case that AI-generated images can receive copyright protection, provided the user made sufficient intellectual contribution. This global legal fragmentation means that AI providers' "ownership transfer" promises may carry entirely different legal weight in different regions, and just how much legal substance these "ownership transfer" promises actually carry remains very much open to question.
Conclusion: Contracts Are the Real Source of Binding Force
Returning to the original question—"I own Claude's output, so why can't I use it to train my own model?" The answer is this: the terms of service you agreed to use contract law, independent of copyright, to restrict specific uses of the output. Ownership gives you the freedom to use it, but the contract claws back the most threatening part of that freedom.
For developers, the pragmatic approach is: carefully read the terms of whatever AI service you're using, and clearly determine whether your business scenario crosses the "training competing models" red line. In an era where AI technology and legal frameworks have yet to fully align, understanding the logic behind these rules matters far more than blindly believing "I own it."
Related articles

Zero-Dependency AI Memory Layer: Agent Memory Without a Vector Database
Explore zero-dependency AI Agent memory layers that work without vector databases. Compare with traditional RAG architectures and learn when lightweight alternatives make more sense.

The Linear Startup Story: From Leaving Coinbase to Redefining Developer Tools
How Linear co-founder Jori Lallo left Coinbase in 2018 to build a developer-first project management tool, defying skeptics to carve out success in a market dominated by Jira, Asana, and Trello.

Why Is AWS S3 Called the Eighth Wonder of the World? The Invisible Power of Cloud Storage
A viral tweet listed AWS S3 as the Eighth Wonder of the World. Explore how S3's eleven 9s durability and architectural ubiquity make it the invisible cornerstone of modern digital civilization.