Google Accused of Illegally Retaining User Data: The Deeper Battle Behind Personal Privacy Lawsuits

A personal lawsuit against Google for illegal data retention highlights the growing battle over user data sovereignty.
A Hacker News user has announced legal action against Google for allegedly retaining customer data illegally. This article examines what constitutes illegal data retention under GDPR and CCPA, reviews Google's history of privacy penalties, analyzes the challenges and significance of individual privacy lawsuits, and explores how AI training data demands are intensifying the battle over data rights.
Event Overview
Recently, a user posted on Hacker News accusing Google of illegally retaining customer data and announcing legal action against the company. Although the post has generated limited discussion so far (with only single-digit upvotes and comments), it touches on a core issue that has long troubled the tech industry and regulators — the boundaries and compliance of user data retention.
In the digital economy era, data is often called "the new oil." This metaphor was first proposed by British mathematician Clive Humby in 2006, suggesting that data itself is like unrefined crude oil, requiring cleaning, analysis, and modeling to unlock its value. Platform companies like Google, Meta, and Amazon have built a digital advertising market worth hundreds of billions of dollars precisely by exchanging free services for user data, then converting it into precision ad-targeting capabilities. However, unlike oil, data is non-exclusive and replicable — once collected, it's extremely difficult to truly "disappear," which is the technical root cause of the data retention problem. The tension between this business model and user privacy rights is increasingly manifesting as legal disputes. This individual lawsuit against Google is a microcosm of this broader trend.

Why Data Retention Has Become a Sensitive Issue
What Constitutes "Illegal Retention" of Data
So-called "illegal retention of customer data" typically refers to situations where companies continue to store user information under the following circumstances:
- Exceeding necessary time limits: Data is stored far beyond the reasonable period required for its original collection purpose;
- Without explicit consent: Personal information is continuously stored without clear user authorization;
- Violating deletion requests: Users have explicitly requested data deletion, but the company fails to fully comply;
- Exceeding scope of use: Data collected for a specific purpose is used for other undisclosed purposes.
Under the EU's General Data Protection Regulation (GDPR) framework, "Storage Limitation" and the "Right to be Forgotten" are explicit legal principles. GDPR officially took effect on May 25, 2018, and is one of the most influential data protection regulations globally. It establishes seven core principles: lawfulness, purpose limitation, data minimization, accuracy, storage limitation, integrity and confidentiality, and accountability. The "storage limitation" principle specifically requires that personal data be kept in a form that permits identification of data subjects for no longer than is necessary for the purposes for which the data is processed. Companies that violate GDPR may face fines of up to 4% of global annual turnover or €20 million (whichever is higher). The "Right to be Forgotten" originated from the landmark 2014 European Court of Justice ruling in Google Spain v. AEPD, which allows individuals under certain conditions to request search engines to remove outdated or irrelevant links related to them. Companies must be able to demonstrate a legitimate basis for their data retention or face substantial fines.
Google's Privacy Compliance Pressure
This is far from the first time Google has been thrust into the spotlight over data handling issues. In recent years, the tech giant has faced privacy-related investigations and penalties across multiple jurisdictions worldwide:
- European regulators have issued fines ranging from tens of millions to hundreds of millions of euros for its data processing practices;
- In a class action lawsuit alleging that "Incognito Mode" still tracked users, Google ultimately agreed to destroy billions of browsing records. This case (Brown v. Google) was filed in 2020 in a U.S. federal court in California, with plaintiffs alleging that Google continued to track user browsing activity through tools like Google Analytics and Google Ad Manager even in Chrome's Incognito Mode, seeking $5 billion in damages. After reaching a settlement in April 2024, Google agreed to delete or de-identify the relevant data and more clearly disclose on the Incognito Mode launch page what data it still collects. This case highlighted the enormous gap between users' reasonable expectations of "privacy" and companies' actual data practices;
- The collection and retention of location data has also repeatedly become a focus of litigation.
These precedents indicate that even lawsuits initiated by individuals may touch on real compliance gray areas in Google's data practices.
The Significance and Challenges of Individual Privacy Lawsuits
David vs. Goliath or a Spark That Starts a Prairie Fire
Facing a tech giant worth hundreds of billions of dollars with a massive legal team, an individual filing a lawsuit undoubtedly faces enormous power asymmetry. The plaintiff needs to:
- Burden of proof: Prove that Google indeed illegally retained their data, which is technically difficult as internal evidence is often inaccessible;
- Legal basis: Clearly cite applicable privacy regulations (such as GDPR, CCPA, etc.). It's worth noting that the United States still lacks a unified federal privacy law, with varying legislative progress across states creating a "patchwork" regulatory landscape. The California Consumer Privacy Act (CCPA) took effect in 2020 as America's first comprehensive state-level consumer privacy law, granting California residents the right to know, the right to delete, and the right to opt out of data sales. In January 2023, its upgraded version, the California Privacy Rights Act (CPRA), further introduced data minimization principles and storage limitation requirements, and established the California Privacy Protection Agency (CPPA) dedicated to enforcement. This fragmented legal environment both increases compliance costs for businesses and creates legal uncertainty for users seeking to assert their rights;
- Damage determination: Prove that data retention caused actual or potential harm.
However, the value of individual lawsuits should not be underestimated. Many far-reaching privacy precedents were driven by individual advocates. For example, Austrian activist Max Schrems' lawsuit against Facebook ultimately overturned the "Safe Harbor" and "Privacy Shield" data transfer frameworks between Europe and the US, profoundly changing the rules governing transatlantic data flows. Schrems filed a complaint with the Irish Data Protection Commission in 2013 regarding Facebook's transfer of European user data to the United States, triggering two landmark European Court of Justice rulings: the 2015 "Schrems I" ruling overturned the 15-year-old Safe Harbor agreement; the 2020 "Schrems II" ruling overturned the replacement Privacy Shield framework, reasoning that US mass surveillance programs could not provide European citizen data with protection equivalent to GDPR. These two rulings forced thousands of companies to reassess their transatlantic data transfer mechanisms and facilitated the signing of the new EU-U.S. Data Privacy Framework in July 2023. Schrems himself founded the non-profit organization NOYB ("None of Your Business"), which continues to actively pursue privacy litigation, having filed hundreds of complaints across multiple EU countries.
From Individual Cases to Collective Action
Once a single user's lawsuit gains attention, it can often evolve into a class action, pooling the power of more affected individuals and creating substantive pressure on companies. Class action lawsuits are a procedure in common law systems that allows one or a few plaintiffs to file suit on behalf of a group of people with the same interests. They hold unique value in the privacy domain: the direct economic loss suffered by a single user from a data breach may be negligible, insufficient to support the cost of independent litigation, but when thousands of affected users come together, the cumulative damages are enough to create a meaningful deterrent for companies. When courts decide whether to "certify" a class action, they typically examine whether the case has commonality, typicality, and adequate representation. In recent years, the EU has also been promoting similar mechanisms — the EU Representative Actions Directive, which took effect in 2023, allows consumer organizations to file cross-border lawsuits on behalf of user groups, further strengthening collective advocacy capabilities. Tech communities like Hacker News serve as important channels for such advocacy actions to gain early exposure.
Data Protection Lessons for Ordinary Users
This incident reminds every internet user to pay attention to their data rights:
- Make good use of data management tools: Platforms like Google provide activity records, data download, and deletion portals — users should review these regularly;
- Understand local privacy laws: Regulations on data retention vary greatly across regions, with EU users typically enjoying more comprehensive rights than those in other areas;
- Maintain evidence awareness: If you suspect a company is handling your data improperly, take screenshots and record timelines promptly to preserve evidence for potential legal action.
Conclusion: The Battle in the Age of Data Sovereignty
Regardless of the final outcome of this individual lawsuit, it reflects an irreversible trend: users' awareness of their own data sovereignty is awakening. As AI large models' hunger for training data grows, the collection, retention, and use of data will become the primary battleground for legal and ethical disputes in the future. Currently, training large language models like ChatGPT, Gemini, and Claude requires massive amounts of internet text, books, news articles, and user-generated content, but whether the collection of this data obtained effective consent from data subjects and whether it complies with the "purpose limitation" principle is triggering intensive legal action. Between 2023 and 2024, landmark cases have emerged in rapid succession, including The New York Times suing OpenAI and Microsoft for copyright infringement, and multiple authors collectively suing Meta and OpenAI for unauthorized use of their works to train AI. In the context of data retention, a core question surfaces: does a company's long-term retention of user data for the purpose of training AI models constitute illegal retention that "exceeds the original collection purpose"? Italy's Data Protection Authority once temporarily banned ChatGPT over similar concerns, requiring OpenAI to provide a legal basis for processing user data. This battle over AI and data rights will profoundly shape the technology governance landscape for the next decade.
For tech giants, transparent data practices and strict compliance management are no longer optional — they are necessary conditions for maintaining user trust and mitigating legal risk. For ordinary users, every attempt at advocacy adds another brick to building a healthier digital ecosystem. This battle over data has only just begun.
Related articles

Getting Started in Machine Learning Research: Essential Paper Reading List and Research Internship Application Path
A complete path from zero to research internship for ML beginners, covering essential classic papers (AlexNet, ResNet, Transformer), paper reading methods, reproduction tips, and practical advice for research internship applications.

Claude Code Hands-On Tutorial: Complete Guide from Installation to Automated Development
Complete guide to Claude Code covering environment setup, permission configuration, Go Goals autonomous loops, Skills system, MCP protocol integration, and version control for automated development.

Gemini 3.7 Flash Release and GPT-5.6 Ultra-Fast Mode: AI Open Source Enters the Ecosystem Era
Google releases Gemini 3.7 Flash for coding and Agent optimization while OpenAI launches GPT-5.6 Ultra-Fast mode with 14x speed gains. AI open source shifts from open models to open ecosystems.