Anthropic Reportedly Building Predictive Surveillance System — AI Safety Pioneer Faces Ethical Backlash

Anthropic faces backlash over reports it's building a predictive surveillance system targeting activists.
Anthropic, the AI safety-focused company behind Claude, is under fire after reports that it's building a predictive surveillance system to monitor activists. The controversy highlights a deep tension between the company's safety-first brand and commercial pressures, raising broader questions about AI governance, the dual-use nature of large language models, and whether corporate self-regulation is sufficient to prevent misuse of powerful AI technologies.
A Report That Sparked Heated Debate
Recently, an article published by the progressive American media outlet The American Prospect ignited widespread discussion on platforms like Reddit. The report directly accused AI company Anthropic of building a predictive surveillance system to monitor activists. The claim quickly set off a fresh round of debate within the tech community over AI ethics, privacy boundaries, and the social responsibilities of technology companies.
In the Reddit thread discussing the report, one comment — "What are we doing here, guys?" — captured the complicated feelings of many tech professionals and observers. For a company that positions itself around the core mission of "AI safety," any involvement in surveillance would stand in sharp contrast to its publicly stated values.

Anthropic's Public Identity and the Potential Contradiction
A Brand Built on "Safety First"
Founded by former OpenAI executives, Anthropic has long positioned itself as the most safety- and alignment-focused force in the AI industry. AI alignment — a core concept in AI safety research — refers to ensuring that an AI system's goals, behaviors, and decisions remain consistent with human intentions and values. The problem is difficult because it operates on multiple levels: first, there's "outer alignment," which asks whether a model's optimization objective truly reflects what humans want; then there's "inner alignment," which asks whether the internal representations a model learns during training are consistent with its surface-level behavior — a model might appear safe and compliant during evaluation but exhibit entirely different behavior in real-world deployment (so-called "deceptive alignment"). Anthropic has invested heavily in alignment research, including Mechanistic Interpretability and Model Evaluations, which is a major reason it attracts top safety research talent.
Its Claude model series emphasizes the concept of "Constitutional AI," which attempts to make model behavior more consistent with human ethical norms through built-in value principles. Specifically, Constitutional AI is a training methodology proposed by Anthropic in 2022. Its core idea is to use a set of explicit principles (the "constitution") to guide AI model behavior. Training proceeds in two phases: in the first phase, the model critiques and revises its own outputs based on the constitutional principles (self-supervision); in the second phase, the revised data is used for RLHF (Reinforcement Learning from Human Feedback) training. Compared to traditional pure human annotation approaches, this method aims to reduce reliance on large numbers of human annotators while making AI value constraints more systematic and auditable. This methodology is a key technical differentiator that sets Anthropic apart from competitors like OpenAI, and it serves as the technical cornerstone of its "safety first" brand narrative. Company founder Dario Amodei has also repeatedly warned in public about the risks that advanced AI could pose.
This is precisely why the label "predictive surveillance" feels so jarring. Predictive surveillance typically refers to using massive datasets and machine learning algorithms to proactively identify, flag, or even anticipate the behavior of specific groups. This technology is not a new invention of the AI era — its lineage traces back to predictive policing in the early 2000s. The most representative system was PredPol (now renamed Geolitica), used by the Los Angeles Police Department, which analyzed historical crime data to predict high-crime areas and time periods. However, multiple independent studies have shown that such systems tend to amplify existing racial and socioeconomic biases — because historical crime data inherently reflects imbalances in law enforcement resource allocation. After 2020, several cities including Los Angeles and Santa Cruz announced they would discontinue or restrict such tools. In the era of large language models, the capabilities of predictive surveillance have undergone a qualitative leap: systems are no longer limited to structured crime record data but can directly process social media posts, communication records, public speeches, and other unstructured text to infer intentions and predict behavior of specific individuals or groups. When this technology is directed at "activists," its potential threat to freedom of speech and the right to assembly is plain for all to see.
Information Gaps That Require Cool-Headed Scrutiny
It must be noted that the current report comes from a single media source and still lacks independent cross-verification from multiple parties. The precise meaning of "building a surveillance system" — whether this is a product actively developed by Anthropic, or whether its models are being used by third parties (such as government agencies or security contractors) for such purposes — is prone to interpretive bias in the absence of complete technical details.
Before forming a judgment, readers should distinguish between deliberate corporate strategy and downstream misuse of technology — two fundamentally different scenarios. Yet even if this controversy ultimately turns out to be the latter, it still raises a serious question: when a company releases a powerful general-purpose AI model, what kind of due diligence obligation does it bear regarding downstream use?
AI and Surveillance: An Inescapable Industry Issue
The Inherent Double-Edged Nature of Large Models
Regardless of the specifics of Anthropic's situation, this controversy reflects a structural problem that cannot be avoided in the generative AI era: powerful language and reasoning models inherently possess capabilities for information aggregation, pattern recognition, and automated analysis — which are precisely the core technologies needed to build surveillance systems.
When a model can rapidly read, summarize, and classify massive amounts of text, and extract correlated information from social media and public records, it is often just one application scenario away from becoming a "surveillance tool." This means that virtually all leading AI companies could face similar ethical scrutiny — the neutrality of technical capabilities does not automatically exempt them from moral responsibility when those capabilities are used for controversial purposes.
This controversy reveals a deeper paradox: even if a model achieves "alignment" in the technical sense — following safety guidelines set by its trainers, refusing to directly generate harmful content — the choice of application scenario remains a human decision. Alignment addresses "whether the model acts according to human intent," but it cannot answer "whether the human intent itself is legitimate."
The Tension Between Commercial Pressure and Founding Mission
AI companies universally face enormous compute costs and fundraising pressure. Understanding this requires knowledge of the cost structure of frontier AI: training a top-tier large language model has surged from tens of millions of dollars in 2023 to hundreds of millions or even billions of dollars in 2025 — and that doesn't include ongoing inference costs (the compute consumed each time the deployed model responds to a user query). Anthropic has raised over $10 billion in cumulative funding to date, with investors including tech giants like Google and Salesforce as well as multiple sovereign wealth funds. Such massive capital investment means the company faces intense pressure to generate revenue.
Government and security sectors tend to be well-funded clients with stable demand. The U.S. Department of Defense's annual IT and AI-related budget runs into the tens of billions of dollars, while consumer-facing AI assistant subscription revenue, though growing rapidly, offers far lower profit margins than government contracts. This economic reality forms the underlying logic of how AI companies weigh "mission" against "survival." For any AI company pursuing commercial sustainability, drawing red lines between revenue growth and founding mission is a very real challenge.
This is also the deeper reason this controversy resonated so broadly — it touches on the conflict between business models and values across the entire AI industry.
The Tech Community's Response: A Mix of Disappointment and Vigilance
Judging by the reactions in Reddit discussion threads, tech professionals displayed a clear mix of disappointment and vigilance. Many pointed out that if even a company flying the "safety" banner could potentially be involved in surveillance, then the reliability of the entire industry's self-regulatory mechanisms deserves a giant question mark.
Behind this reaction lies a more pervasive anxiety: AI governance currently relies heavily on corporate self-regulation, lacking a robust external regulatory framework. Globally, AI regulation presents a fragmented landscape. The EU's AI Act, which officially took effect in 2024, is the world's first comprehensive AI legislation, classifying "real-time remote biometric identification" and "social scoring systems" as "unacceptable risk" categories to be prohibited, while imposing strict restrictions on predictive policing tools. By contrast, the United States has taken a softer regulatory approach centered on executive orders and industry self-regulation — while the Biden administration's 2023 AI executive order introduced safety assessment requirements, it lacked mandatory enforcement mechanisms, and the Trump administration has since revoked the order. China, through regulations such as the Interim Measures for the Management of Generative AI Services, requires AI service providers to bear responsibility for output content. Against this backdrop of uneven regulation, the constraints AI companies face vary enormously across different jurisdictions, making the reliability of corporate self-regulation a critical variable.
When a gap appears between a company's public commitments and its actual business practices, the public has virtually no effective means of verification or accountability. This reminds us that AI ethics cannot remain at the level of corporate mission statements — it requires transparent disclosure mechanisms, independent audits, and enforceable legal standards.
Conclusion: What Exactly Are We Building?
"What are we doing here, guys?" — this simple question may be the one the entire AI industry most urgently needs to take seriously right now. Technology itself has no agenda, but how technology is used, whom it serves, and where boundaries are drawn are choices every AI company must answer publicly.
While we await more facts to surface, this controversy offers at least one important warning: AI safety is not only about whether a model might "go rogue" — it's equally about whether it might be used to suppress fundamental human rights. Truly responsible AI development requires not just technical alignment, but transparency and steadfastness in values.
Related articles

Production-Grade AI Agent State Verification: Four Mainstream Strategies and Risk-Tiered Best Practices
Explore four key strategies for post-operation state verification in AI Agent workflows — trust, read-back, idempotency, and monitoring — with risk-tiered production guidance.

Taming AI Context Bloat with Kanban: A Practical Guide to Parallel Agent Workflows
A deep dive into a Kanban-based solution for AI context bloat, using structured external memory, parallel Agent workspaces, and scripted lifecycle management.

Claude Code Installation Guide for Beginners: From Environment Setup to AI-Built Games
Complete beginner's guide to installing Claude Code, covering Node.js, Python, Git setup, CC Switch model configuration, and building a Minesweeper game deployed to GitHub Pages.