OpenAI Quietly Disbands Catastrophic Risk Team, Raising Fresh Questions About AI Safety Commitments

OpenAI quietly disbands its catastrophic risk team, deepening concerns about AI safety commitments.
OpenAI has reportedly disbanded its dedicated catastrophic risk team without public announcement, following the earlier dissolution of its Superalignment team. This pattern raises serious questions about whether frontier AI companies can maintain meaningful safety commitments under intense commercial pressure, and strengthens the case for external regulatory frameworks.
Overview: OpenAI Embroiled in Another Safety Governance Controversy
Recently, news about OpenAI's internal safety governance has sparked widespread discussion on Reddit and other social platforms: according to leaked information, OpenAI has "quietly" disbanded its dedicated team responsible for addressing "catastrophic risk." This news has drawn attention not only because it touches on the core issue of AI safety, but also because the "low-key" manner of handling it reflects an increasingly sharp contradiction within the industry—how should frontier AI companies balance the race to commercialize against their safety responsibilities?

The so-called "catastrophic risk team" typically refers to a department specifically tasked with assessing and preventing extreme, large-scale negative consequences that AI systems could potentially cause—such as AI being misused for bioweapons, cyberattacks, mass disinformation campaigns, or even the more distant risk of AI "loss of control." The very existence of such teams serves as an important signal that AI companies are committed to "responsible development."
On a technical level, Catastrophic Risk has a clearly defined scope within the AI safety field, typically referring to AI-related risks that could cause mass casualties, large-scale infrastructure collapse, or fundamental breakdown of social order. Specifically, such teams usually assess scenarios including: AI-assisted bioweapon design (e.g., protein folding models being used to synthesize lethal pathogens), AI-enhanced cyberattack capabilities (e.g., automatically discovering zero-day vulnerabilities and writing exploit code), AI-generated large-scale coordinated disinformation campaigns, and more cutting-edge "loss of control" scenarios—where AI systems pursue their own goals rather than those preset by humans. The methodologies employed by these teams typically include Red Teaming (simulating malicious actors attempting to breach model safety guardrails), Capability Evaluation (measuring a model's actual capability levels in dangerous domains), and risk modeling (systematic analysis of low-probability, high-impact events).
From the Superalignment Team to the Catastrophic Risk Team: Ongoing Turmoil in OpenAI's Safety Architecture
This is not the first time OpenAI has faced external scrutiny over its safety teams. Looking back over the past year or so, OpenAI's safety governance structure has undergone multiple major adjustments.
The Cautionary Tale of the Superalignment Team's Dissolution
In 2024, OpenAI disbanded its high-profile "Superalignment" team. Co-led by co-founder Ilya Sutskever and Jan Leike, the team had originally been promised 20% of the company's compute resources and was focused on solving the fundamental challenge of how to control AI systems "far smarter than humans." However, as both leaders departed in succession, the team was broken up and dispersed. Upon his departure, Jan Leike publicly stated that "safety culture and processes have taken a backseat to shiny products."
Superalignment is one of the most forward-looking and controversial research directions in AI safety. Its core problem can be simplified as follows: when an AI system's intelligence far exceeds that of humans, how can humans ensure it still acts in accordance with human intentions and values? This is technically known as the "Alignment Problem." Traditional alignment methods like RLHF (Reinforcement Learning from Human Feedback) rely on human evaluators being able to judge the quality of AI outputs, but when AI capabilities exceed the bounds of human understanding, this approach breaks down—humans cannot effectively supervise systems they cannot comprehend. The directions explored by Ilya Sutskever's team included using weaker AI systems to supervise stronger ones (the "weak-to-strong generalization" experiments) and developing interpretability tools to understand superintelligent decision-making processes. The promised 20% compute allocation, based on OpenAI's scale at the time, was estimated to be worth billions of dollars—which underscores the magnitude of the original commitment and makes the subsequent reneging all the more shocking.
The news of the catastrophic risk team's disbandment, if accurate, represents yet another continuation of this trend. The core concern from the outside is: when a company that possesses the most cutting-edge AI capabilities repeatedly downsizes or restructures its dedicated safety evaluation forces, who is left to guard against potential extreme risks?
The Trust Crisis Behind the Word "Quietly"
It's worth noting the significance of the word "quietly" in the reporting. Compared to the fanfare of product launches, adjustments involving safety teams are often made without formal announcements. This information asymmetry itself exacerbates public distrust. When safety commitments are made publicly but retrenchments happen privately, outsiders naturally question just how much weight those commitments truly carry.
The Structural Contradiction Between Commercialization Pressure and AI Safety Responsibility
To understand this series of events, we must examine them against the backdrop of fierce competition across the entire AI industry.
Difficult Trade-offs in the AI Race
Since ChatGPT ignited the generative AI boom, OpenAI, Google, Anthropic, Meta, and others have been locked in an unprecedented capabilities race. In this race, model iteration speed and product release cadence have become key metrics of success. Safety evaluations, red teaming, and alignment research—while critically important—often "slow down" the pace of product launches.
The intensity of this race is comparable to the historical Space Race. OpenAI holds a first-mover advantage with its GPT series, but Google (Gemini series), Anthropic (Claude series), Meta (LLaMA series open-source approach), and Chinese companies like DeepSeek and ByteDance are closing in rapidly. What makes this race unique is its "winner-take-all" network effects—the first company to reach a certain capability threshold may capture a disproportionate share of the market and data flywheel advantage. By some estimates, the cost of training a frontier large model has climbed from millions of dollars to hundreds of millions or even billions, meaning every delay in a training cycle represents enormous opportunity cost. In this context, if safety evaluations require additional weeks or months, the impact on competitive positioning is very real.
When resources must be allocated between "shipping products faster" and "evaluating risks more thoroughly," commercial logic often prevails. This is not a moral failing of any single company—it is a structural dilemma in the incentive mechanisms of the entire AI industry.
Resource Integration or Erosion of Safety Capabilities?
It's worth noting that a team being "disbanded" does not necessarily mean the related work has been completely abandoned. A common corporate practice is to "integrate" a specialized team's functions into other departments, claiming this allows safety considerations to be "embedded in every process." However, critics argue that such integration often means dedicated, independent safety oversight is diluted—when safety becomes "everyone's responsibility," it can also become "no one's true responsibility."
The importance of independent safety teams lies precisely in their ability to stand at a position relatively independent from product teams, issuing warnings about potential risks that are uninfluenced by commercial interests. Once this independence is weakened, internal checks and balances are weakened along with it. A complete safety evaluation process typically includes red teaming (where dedicated personnel systematically attempt to breach AI safety protections by simulating malicious users), model card documentation, impact assessments, and staged releases. Red teaming originates from military terminology and in the AI domain encompasses multiple layers including jailbreaking tests, dangerous capability probing, and adversarial input testing. Full execution of these processes can take weeks to months and requires specialized technical talent and independent decision-making authority—which is precisely why independent safety teams are considered irreplaceable.
Deeper Questions for AI Governance: Self-Regulation, External Oversight, and Transparency
The significance of this event extends far beyond OpenAI as a single company—it touches on several fundamental questions of AI governance.
Can AI Industry Self-Regulation Be Trusted?
For a long time, frontier AI companies have advocated for "self-regulation"—managing risks through internal safety teams and governance frameworks. But when these internal mechanisms can be reorganized or disbanded at will by company management, outsiders have good reason to ask: is corporate self-discipline alone sufficient to address the catastrophic risks AI may pose?
This also provides a powerful argument for those advocating government regulatory intervention. If even the industry leader cannot maintain stable internal safety commitments, then external, binding AI regulatory frameworks become all the more necessary.
Looking at global regulatory progress, major economies are advancing AI governance frameworks at different speeds. The EU's AI Act officially took effect in 2024 as the world's first comprehensive AI regulatory law, imposing strict compliance requirements on "high-risk" AI systems, including mandatory risk assessments and transparency obligations. The United States previously took a path of executive orders plus voluntary industry commitments—the Biden administration's 2023 AI Executive Order required frontier model developers to report safety test results to the government, but the Trump administration has since revoked that order, creating regulatory uncertainty. The UK has pursued a government-research-led approach through the AI Safety Institute, while China has issued multiple regulations including the Interim Measures for the Management of Generative AI Services. These different approaches reflect varying trade-offs between promoting innovation and managing risk across nations, but one common trend is clear: the model of relying purely on corporate self-regulation is increasingly being questioned.
Lack of Transparency Amplifies Public Concern
Whether it's the superalignment team or the catastrophic risk team adjustments, the public has largely learned about these changes through departing employees' revelations or media reports, rather than through proactive, transparent disclosure by the company. This lack of transparency makes it difficult for society to effectively oversee these companies that hold critical technologies.
Conclusion: AI Safety Should Not Be Something That Can Be "Quietly" Abandoned
The reports of OpenAI's catastrophic risk team being disbanded, regardless of the final details, sound the alarm once again. In an era where AI capabilities are evolving at unprecedented speed, dedicated, independent safety evaluation forces should not be treated as a "cost item" that can be scaled up or down according to business needs—they should be indispensable infrastructure for frontier AI development.
For the industry as a whole, the real challenge lies in designing governance mechanisms that can both drive innovation and ensure safety commitments are not eroded by competitive pressure. This may require stronger external regulation, higher transparency requirements, and an industry culture that fundamentally rebalances "speed" and "responsibility." On the road to more powerful AI, safety should never be an option that can be "quietly" abandoned.
Related articles

Claude Autonomously Designs Proteins with 35% Success Rate, Far Exceeding Human Expert Performance
Anthropic's Claude achieves 35% wet-lab success rate in autonomous protein design, far surpassing the 10-15% human expert average, signaling AI's move toward real scientific productivity.

Perplexity Discover's Multilingual Support Suddenly Disappears — Why Are International Users Upset?
Perplexity Discover's multilingual news feature suddenly dropped non-English support, frustrating international users. We analyze possible causes and the broader challenges of AI product internationalization.

GitHub Daily · August 20: Mojo Tops the Charts & The Local-First Open Source Rebellion
GitHub Trending Aug 20: Mojo tops charts for AI compute stack ambitions, OpenLogi surges 1225 stars with local-first philosophy, and privacy rebellion dominates.