Sony and Warner Sue Anthropic: Three Ways the AI Copyright Battle Could End

Sony and Warner sue Anthropic over AI training data, threatening fines, licensing mandates, or model retraining.
Sony Music Publishing and Warner Chappell have sued Anthropic, alleging mass use of pirated content to train Claude. The case could result in fines (likely negligible for a $60B company), licensing fees (a more sustainable model), or the nuclear option of retraining models from scratch. The outcome could set a landmark precedent affecting every AI company's approach to training data and reshape copyright law for the AI era.
The Story: Two Music Giants Take on Anthropic
Sony Music Publishing and Warner Chappell have filed a lawsuit against AI company Anthropic, alleging that it used tens of thousands of pirated works in training its large language model, Claude. At the heart of the complaint is the accusation that Anthropic obtained copyrighted content for model training through mass torrenting, web scraping, and bulk downloading.
Torrenting is a peer-to-peer file-sharing technology based on the BitTorrent protocol, where users simultaneously download file fragments from multiple nodes and reassemble them — a method commonly used to distribute large files. In the context of AI training, "mass torrenting" means systematically acquiring vast amounts of copyrighted content from torrent networks. Web scraping, on the other hand, involves using automated programs (crawlers) to extract data from websites in bulk. AI companies typically deploy large-scale crawler clusters that traverse millions of web pages in a short time to extract text, code, and other content. These techniques are not inherently illegal, but when used to obtain copyrighted works at scale without authorization, they enter a legal gray area.
Anthropic was founded in 2021 by former OpenAI executives Dario Amodei and Daniela Amodei, with "AI safety" as its core philosophy. As of 2024, Anthropic's valuation has surpassed $60 billion, backed by massive investments from tech giants including Amazon (up to $12.5 billion) and Google. Its flagship Claude series of large language models is widely regarded as one of GPT-4's primary competitors, particularly excelling in long-context understanding, code generation, and safety alignment. Anthropic has consistently positioned itself as a champion of "responsible AI development," which makes this copyright lawsuit particularly damaging to its brand image.
In response to these allegations, Anthropic has denied the claims and stated it will vigorously defend itself. While this lawsuit may seem like just another AI copyright dispute, its potential ramifications extend far beyond a typical damages case — it could strike at the very foundations of how the entire AI industry operates.

The Core Dispute: Fines, Licensing, or Retraining?
What makes this AI copyright infringement case so widely discussed is that different rulings would have vastly different impacts on the AI industry. There are currently three main remedies on the table.
Fines: Potentially Just a "Cost of Doing Business"
For an AI unicorn like Anthropic, valued at tens of billions of dollars, a simple financial penalty would likely be a drop in the bucket. As one community member pointed out: "Fines would probably just become a cost of doing business."
This is precisely what many critics fear. If the price of copyright infringement is merely a predictable expense, large AI companies have every incentive to act first and ask forgiveness later — training powerful models on massive datasets to capture market share, then paying fines after the fact. This model effectively internalizes the cost of breaking the law and paradoxically incentivizes disregard for copyright.
Licensing Fees: A More Sustainable Path
Compared to one-time fines, establishing a content licensing fee mechanism is widely considered more aligned with the industry's long-term health. This would mean AI companies must reach agreements with copyright holders and pay reasonable fees for the use of training data.
This model is already taking shape in some areas. For example, OpenAI has signed content licensing agreements with several news publishers, attempting to shift from a "infringe first, compensate later" approach to a "license first, use later" virtuous cycle. For music rights holders, a sustainable licensing framework may prove more valuable in the long run than a one-time payout.
Retraining the Model: A Nuclear Option for the Industry
The most aggressive and deterrent option is requiring Anthropic to destroy or retrain the model from scratch. If this demand is upheld by the court, it would send seismic shockwaves through the entire AI industry.
Retraining a large language model doesn't just mean tens or even hundreds of millions of dollars in compute costs gone to waste — it also means having to rebuild a completely "clean" training dataset. To understand how extreme this demand is, consider the scale of current large model training: Meta's Llama 3, for instance, used over 15 trillion tokens of training data for its 405B parameter version, consumed approximately 30 million GPU hours of compute, and cost potentially hundreds of millions of dollars. The more critical technical challenge is that once a model is trained, the "knowledge" from copyrighted content is encoded in a distributed fashion across billions or even trillions of weight parameters — you can't simply "delete" the influence of specific works. This is the fundamental reason why the "retrain" demand is so extreme from a technical standpoint.
Given that virtually all mainstream large models rely to some degree on massive web-crawled data, establishing this precedent could force every AI company to re-examine the legality of its data sources.
The Deeper Dilemma: The "Original Sin" of AI Training Data
This case reflects a "data original sin" problem that the entire generative AI industry cannot avoid.
Nearly every major language model's capabilities are built on learning from massive amounts of internet data, which inevitably includes vast quantities of copyrighted works — books, articles, song lyrics, code, and more. AI companies generally invoke the "fair use" doctrine in their defense, arguing that the way models learn from data is analogous to how humans read and learn, constituting transformative use.
Fair use is a legal principle established under Section 107 of U.S. copyright law that permits the unauthorized use of copyrighted works under certain conditions. Courts typically evaluate four factors: the purpose and character of the use (whether it is commercial and whether it is transformative), the nature of the copyrighted work, the amount and substantiality of the portion used, and the effect on the original work's market value. AI companies argue that their training process constitutes "transformative use" — the model doesn't copy the original works but rather learns language patterns and knowledge from them to generate entirely new content. Critics counter that when AI models can reproduce copyrighted training content nearly verbatim — such as complete song lyrics or book passages — the "transformative" argument is significantly weakened. From 2023 to the present, U.S. courts have not issued a definitive ruling on this issue across multiple AI copyright cases, leaving the legal landscape fraught with uncertainty.
Copyright holders don't buy the AI companies' logic. They argue that AI companies used creators' work at massive scale without authorization or compensation to build immensely valuable commercial products — essentially free-riding on creators' labor.
Interestingly, Sony and Warner, as two of the music industry's biggest players, have historically taken an aggressive stance on copyright enforcement. Sony Music Publishing is the world's largest music publisher, holding the rights to works by legendary artists like The Beatles and Bob Dylan; Warner Chappell is the world's third-largest music publisher, representing a vast catalog of classic songs from artists including Led Zeppelin and Madonna. Both companies have accumulated extensive copyright litigation experience in the digital music era — from the early battle against Napster to licensing negotiations with streaming platforms like Spotify. Their decision to target Anthropic likely reflects a strategy to use the judicial system to establish industry standards for music copyright pricing in the AI era.
Industry Watch: The Ripple Effect of Legal Precedent
Regardless of the final outcome, this case could become a landmark precedent in the AI copyright space. Notably, the Anthropic case is not an isolated event but an important chapter in the wave of AI copyright litigation over the past two years. The New York Times sued OpenAI and Microsoft in late 2023, alleging unauthorized use of news content for model training; Getty Images sued Stability AI, claiming its image generation model Stable Diffusion infringed on the copyrights of millions of photos; a group of visual artists filed a class-action lawsuit against Stability AI, Midjourney, and DeviantArt; and in the code domain, GitHub Copilot also faces a copyright class-action suit from developers. From text to images, from music to code, virtually every creative domain is grappling with copyright issues around AI training data, and the rulings in these cases will inform one another, collectively shaping the copyright legal framework of the AI era.
If courts side with copyright holders — especially if they support extreme demands like retraining or destroying models — the entire AI industry will be forced into a new phase where data compliance costs skyrocket. Smaller AI companies may be eliminated by their inability to afford licensing fees, further consolidating the industry around giants with data resources and financial muscle.
If courts uphold AI companies' "fair use" arguments, it could further erode content creators' bargaining power and intensify the standoff between creative industries and tech companies.
The most realistic outcome is likely a settlement — the establishment of some form of licensing payment mechanism. This would allow copyright holders to receive financial compensation while giving AI companies legal legitimacy for their data usage, representing a relatively balanced solution.
Copyright Rules Are Being Rewritten for the AI Era
The lawsuit facing Anthropic is, at its core, a head-on collision between the legacy copyright framework and emerging AI technology. The questions it raises go far beyond the fate of a single company — they concern the operating rules for an entire industry going forward: What price should AI companies pay for using creative content to train their models?
In an era of breakneck technological advancement, the lag between legal and ethical frameworks and innovation inevitably breeds vast gray areas. This case may be precisely the catalyst society needs to confront this issue head-on and establish clearer rules. Regardless of which side the outcome favors, one thing is certain — the era of "free lunches" for AI training data is coming to an end.
Related articles

Glasp Firefox Extension: A Detailed Guide to Free AI Highlighting & Smart Summarization
Glasp launches on Firefox with multi-color highlighting for web pages, PDFs, and YouTube videos, AI summaries via ChatGPT, Claude & Gemini, plus free export to Notion and Obsidian.

Wealthfolio: A Local-First Open-Source Personal Finance Tool
Wealthfolio is an open-source, local-first personal finance app for investment tracking, net worth, and expense management — with no accounts, no subscriptions, and full data privacy.

Gojo: Turn Your MacBook's Notch into a Voice Input and Productivity Hub
Gojo is an open-source tool that transforms the MacBook notch into a feature panel with local voice dictation, clipboard history, window controls, and more — all processed locally for privacy.