Former Anthropic Researcher's Resignation Tweet Goes Viral: The Truth Behind AI Safety's 'Conspiracy Theory'

A viral resignation tweet exposes the tangled web of money, ideology, and power behind AI safety governance.
Former OpenAI and Anthropic pretraining researcher Jacob Coxon's resignation statement became one of tech's most viral tweets at 172M impressions, drawing responses from Sam Altman, Elon Musk, Dario Amodei, and Bernie Sanders. But the episode's deeper story involves suspicious timing — the WSJ article ran 18 minutes before the tweet — alongside Dario's proposed "embedded evaluators" framework naming Meter as overseer, and a densely overlapping network of EA-aligned investors and employees connecting Meter to Anthropic. The analyst's verdict: Dario's intentions to shape AI's direction are real, but a "grand cabal" is unproven, and Altman and Musk's endorsements may be strategic rope-a-dope.
A single tweet racking up 172 million impressions thrust the AI safety debate into an unprecedented spotlight. Jacob Coxon's resignation statement — posted by a former pretraining researcher at both OpenAI and Anthropic — not only drew responses from heavyweights like Bernie Sanders, Sam Altman, Elon Musk, and Dario Amodei, but also ignited a sweeping debate about whether AI safety is really just a power game.
The Tweet That Rewrote Tech Twitter History
In his resignation statement, Jacob Coxon wrote: "I resigned from Anthropic today. For the past three years I've done pretraining research at OpenAI and Anthropic. Neither company is being responsible. They are racing toward a self-improving superintelligence and gambling with all of our lives."
The post's spread was staggering — roughly 800,000 likes, 200,000 reposts, and over 170 million impressions. The video creator behind the analysis stated plainly that this was the most widely circulated tweet he had ever seen in the entire tech space — unprecedented in AI safety circles or broader tech Twitter. Even Elon Musk marveled that he had never seen an account with virtually no prior engagement or follower base go viral like this.
Several alignment research leads within Anthropic publicly voiced support. Evan, a head of alignment science, retweeted and stated: "I personally think the probability that AI kills all of humanity within the next decade is greater than 10%… We don't currently have a solution to the superintelligence alignment problem, and we're not obviously on track to get one." That statement laid bare the anxiety about safety risks simmering inside the company.

Suspicious Coincidences on the Timeline
The analyst, while describing himself as "not really a conspiracy theorist," nonetheless identified a string of intriguing coincidences.
First, the media timing. The Wall Street Journal published an article citing Coxon just 18 minutes before he posted his viral tweet. In other words, a researcher about to resign had already granted exclusive access to a journalist before going public. As the analyst put it: "I've never heard of someone wanting to quit their job, then contacting media to prepare to blast their former employer."
Adding to the suspicion: a Netflix documentary about AI dangers — reportedly three years in the making — launched that same week; Time magazine's cover story happened to be "Pause AI." And Coxon's account itself raised eyebrows — created this year, renamed twice, with just two posts total: the original resignation statement and an AMA roughly seven days later.

These individually innocuous facts, stacked together, created fertile ground for questions about whether the whole thing was carefully orchestrated.
The Central Accusation: Dario's "Evaluator" Chess Move
The real ignition point came from an article Dario Amodei published shortly after, titled We Must Pace the Frontier. He laid out a three-step plan, and the most controversial piece was the first: embedded evaluators.
Dario argued that every frontier AI company should grant third-party evaluation teams continuous access with "near-employee-level" permissions to verify safety practices — covering not just finished models but the training process itself. The only evaluation body he named was Meter.
Critics extrapolated a damning logic chain: if Meter holds the power to determine which labs are "qualified" to pursue superintelligence, and Meter has deep ties to Anthropic and Dario himself, then so-called "independent third-party oversight" could become a backdoor mechanism for industry control.
Noteably, Sam Altman — who doesn't typically align with Dario — along with Elon Musk both publicly endorsed Dario's "slow down" argument. When policymakers and rival CEOs all chime in simultaneously, seeds of suspicion plant themselves naturally.

Embedded Evaluators is a regulatory model proposed by Dario in that piece — distinct from traditional "post-hoc review." It would require third-party organizations to have continuous, near-employee-level internal access during the model training phase, including codebases, experiment logs, and training pipelines. The rationale: waiting until a model is fully trained before conducting safety evaluations is often too late, since dangerous capabilities may emerge mid-training. The controversy, however, stems precisely from this "deep embedding" — the evaluating body would effectively need to be permanently stationed inside the lab, gaining access not just to safety data but potentially to core technical details and trade secrets. Critics argue that once such permissions are granted to a specific institution, that institution gains a de facto veto over the entire industry's development pace, rather than merely serving as an outside observer.
The Effective Altruism Network
The most substantive part of the analysis was the analyst's attempt to map out the overlapping network between Meter, Anthropic investors, and employees — most of whom are deeply tied to the Effective Altruism (EA) community.
Based on publicly available information he compiled:
- Meter was incubated out of ARC (founded by Paul Christiano), and received funding from ARC, the Survival and Flourishing Fund (SFF), Open Philanthropy, and FTX;
- Open Philanthropy is funded through Dustin Moskovitz's Good Ventures; Moskovitz is an early Anthropic investor;
- Open Philanthropy co-founder Holden is married to Dario's sister Daniella, and Holden, Daniella, and Dario once lived together;
- Paul Christiano also once lived with Dario and serves as a trustee of Anthropic's Long-Term Benefit Trust (LTBT);
- Paul's wife Ajaya works at Meter and previously handled AI safety at Open Philanthropy;
- Even Sam Bankman-Fried, the imprisoned FTX founder, was an early Anthropic investor — and FTX provided Meter with over a million dollars in funding.
Piecing these nodes together, the skeptics' argument is that the supposedly "independent third-party evaluation" is deeply entangled with Anthropic's investors and employees.

Investor David Sacks's critique is cited as a representative view: "Stop pretending Meter is independent when it is intertwined with Anthropic's investors and employees. Stop pretending you need an antitrust exemption to form a cartel." Dario did indeed suggest in his piece that AI companies should coordinate on how AI gets deployed — which critics read as a "legitimized monopoly coalition."
Effective Altruism (EA) is a philosophical and social movement centered on rigorous quantification and utilitarian reasoning, advocating for the allocation of resources to wherever they can do the most measurable good. In the AI safety space, the EA community has long treated "preventing superintelligence from destroying humanity" as its top priority, channeling substantial funding through foundations like Open Philanthropy and the Survival and Flourishing Fund into AI safety research organizations, policy advocacy groups, and even some AI labs themselves. The community is characterized by an unusually tight-knit network — core members often graduated from the same elite universities, share the same existential risk narrative framework, and are deeply bound together through cross-investment, co-founding, marriages, and cohabitation. Critics point out that when "regulator," "regulated entity," and "funder" all overlap within the same ideological circle, independence becomes a hollow label — not because of malicious collusion, but because the cognitive frameworks and incentive structures are simply inseparable.
The Analyst's Actual Verdict
After laying out all the conspiracy-friendly material, the analyst offered a relatively sober conclusion.
His assessment: Dario wanting to shape the direction of AI development, push for regulation, and maintain close ties with Meter — these are essentially facts, ones Dario himself has largely acknowledged. But to call it a "grand cabal" controlling everything is something that cannot be proven.
The analyst also proposed his own "counter-conspiracy": Sam Altman and Elon Musk's public agreement with Dario's "slow down" stance might actually be a form of strategic rope-a-dope — saying "yes, yes, we should slow down" while privately planning to let Dario restrain himself as they "shoot for the moon" at full speed.
The video also ended with a twist: Coxon had claimed on a live television appearance that his resignation was coordinated with no third parties, but Wall Street Journal reporting showed he had been in contact with a policy director at an anti-AI policy organization before resigning. The analyst concluded that Coxon had "told quite a few lies."
The "rope-a-dope" logic the analyst describes is not without historical precedent in tech competition. Publicly endorsing a competitor's self-restraint initiative while privately accelerating is a classic strategy — using the appearance of consensus to mask differentiated competitive positioning. For Sam Altman, verbally supporting a "slowdown" earns goodwill with regulators without requiring OpenAI to meaningfully decelerate. For Elon Musk, endorsing limits on Anthropic doesn't contradict xAI's own full-throttle pursuit. The elegance of this strategy is that it can trap the initiative's originator (Dario) in a moral bind: violate your own rules and lose credibility; follow them and you've unilaterally slowed yourself down.
What This Episode Actually Reveals
Stripping away the entertainment value of the conspiracy framing, this incident reflects a genuine dilemma in AI safety governance: when the key actors in "independent oversight" share investors, ideologies, and even personal relationships with the entities being overseen, the credibility of that oversight is inherently open to question.
Whether Coxon's motivations were purely driven by public responsibility or whether he was leveraged by a larger game, the 172 million impressions have already demonstrated that public concern about superintelligence risk is surging. And whether Dario's proposed "embedded evaluators" framework represents responsible industry self-regulation or a disguised concentration of power — that, arguably, is the question worth continuing to ask beneath all the noise.
Related articles

Jev + Claude Code: How Fast-and-Slow Thinking Can Cut AI Agent Costs by 90%
Jev is a System 1 frontend model that outputs option probabilities in milliseconds — not text. See how it pairs with Claude Code and GPT-6 Astra to cut AI agent costs by up to 90%.

Budget AI Coding Setup: Connecting VSCode + Claude Code to DeepSeek
Step-by-step guide to setting up a budget AI coding environment using VSCode + Claude Code connected to the DeepSeek API. Covers API setup, extension install, config, and verification.

Ginger Cinnamon Tea for Cold Relief: The Physiology Behind a Traditional Herbal Remedy
Antibiotics don't work on colds. Try ginger cinnamon tea instead. Learn the simple recipe and the physiology behind how ginger dilates blood vessels and thins mucus.