Frontier AI Model Safety Evaluation: How Do Governments Make These Decisions?

A deep look at how (and whether) governments actually evaluate the safety of frontier AI models before release.
When frontier AI models are released, a critical question goes unasked: how do governments decide they're safe? This article examines the opacity of current AI safety governance — from voluntary corporate frameworks like OpenAI's Preparedness Framework and Anthropic's RSP, to the limited authority of bodies like the US AI Safety Institute — and explores what transparent, accountable evaluation would actually require.
A Critical Question We're Not Asking
When OpenAI, Anthropic, and other companies release their latest frontier AI models, public attention tends to focus on what these models can do — write code, generate images, perform complex reasoning. Yet a far more fundamental question rarely gets asked: By what standards do governments actually determine that these frontier models are safe to release?
As relevant reporting has noted: "What conversations between governments and Anthropic and OpenAI actually look like remains unclear." That statement reveals a deeply unsettling reality at the heart of AI safety governance — the opacity of the evaluation process.

The Current State of Frontier AI Model Evaluation
Who's Actually Making the Decisions?
In the United States and other leading AI development nations, frontier model safety evaluations typically involve three types of actors: AI companies' own safety teams, relevant government agencies (such as the US AI Safety Institute), and independent third-party evaluators. However, the boundaries of responsibility among these three groups remain far from clear.
The current evaluation model relies heavily on voluntary corporate commitments. Leading companies like OpenAI and Anthropic have each published their own "Responsible Scaling Policies" or "Preparedness Frameworks," pledging to initiate additional safety testing when models reach certain capability thresholds. OpenAI's Preparedness Framework categorizes model risk into four levels — low, medium, high, and critical — covering dimensions such as cyberattacks, bioweapons, and autonomous replication. Anthropic's Responsible Scaling Policy introduces the concept of "AI Safety Levels" (ASL), from ASL-1 through ASL-4, with progressively stricter safety requirements at each tier. Yet both frameworks share a common limitation: the evaluating party is still the company itself. Governments function as information recipients rather than active auditors, and the actual implementation of these commitments is rarely verifiable by outside parties.
The US AI Safety Institute: Advisor or Regulator?
The US AI Safety Institute (AISI), established in November 2023 under the National Institute of Standards and Technology (NIST) framework, is currently the closest thing to an official government evaluation body — with a parallel institution launched in the UK around the same time. However, AISI's evaluation work depends on AI companies voluntarily providing model access; it holds no mandatory review authority, and its findings are not directly tied to any release approval process. This means its role more closely resembles a "technical advisor" than an "approval authority," making it difficult to exert genuine regulatory constraint — an important backdrop for understanding the current regulatory vacuum.
The Problem of Information Opacity
The most fundamental issue is this: what information does a government actually obtain when it approves — or tacitly allows — the release of a frontier model? What standards guide the evaluation? These critical details are largely not made public.
When the decision-making process operates as a black box, the public has no way to assess:
- Whether evaluation standards are sufficiently rigorous
- Whether regulatory bodies possess the technical capacity to understand model risks
- Whether the evaluation data submitted by companies is complete and objective
This opacity not only erodes public trust — it leaves the entire AI safety governance system without the accountability mechanisms it urgently needs.
Why Frontier Model Safety Evaluation Matters
The Potential Risks of Frontier Models
Frontier AI models require dedicated evaluation precisely because they can generate "catastrophic risks" — scenarios including assistance with bioweapon design, large-scale cyberattacks, and automated fraud. As model capabilities continue to leap forward, the likelihood of these risks moving from theoretical to real increases accordingly.
In AI safety research, "catastrophic risk" carries a precise technical meaning. Take biosecurity as an example: the core concern is whether a model can provide non-experts with sufficient "uplift" — meaningfully lowering the knowledge barrier to synthesizing highly pathogenic agents or carrying out a biological attack. Assessments from institutions like RAND and the Johns Hopkins Center for Health Security suggest that today's top models already demonstrate knowledge in certain biochemical domains that exceeds the average college student — though whether this constitutes an operationally actionable threat remains contested. For this reason, safety evaluations should not be a formality. They need to genuinely answer: would widespread deployment of this model significantly enhance the capabilities of malicious actors? Are the existing safety guardrails adequate?
The Challenges Facing Government Regulatory Capacity
A deeper question looms: do governments actually have the technical expertise to evaluate frontier AI models? These systems are far more complex than traditional regulatory subjects, and agencies broadly face shortages of technical talent and lagging specialized knowledge — making independent, in-depth safety assessments extremely difficult.
This creates a classic "information asymmetry" — the regulated party (AI companies) holds vastly more technical knowledge and internal data than the regulator. This asymmetry is particularly extreme in AI: training a frontier model requires hundreds of millions of dollars and thousands of high-end GPUs. Regulatory bodies lack the resources to independently replicate these models, and without company cooperation, deep technical audits are nearly impossible. By contrast, financial regulators can verify data through standardized accounting principles; food and drug regulators can conduct independent laboratory tests. AI regulators currently lack equivalent independent verification tools. This structural deficit makes the risk of "regulatory capture" — where a regulator's stance gradually aligns with the interests of the companies it oversees — especially pronounced in the AI domain. In this environment, regulation risks becoming a passive rubber stamp on corporate self-assessment rather than a genuinely independent safety review.
Toward Transparent AI Safety Governance
Core Mechanisms That Need to Be Built
Breaking through this impasse will require coordinated effort from the industry and policymakers to drive the following improvements:
Make evaluation standards public. Governments and companies should publish the basic frameworks and capability thresholds used in frontier AI safety evaluations, so the public can understand the basis for critical decisions — rather than simply being told a model has "passed a safety review."
Introduce independent third-party evaluation. Entrusting safety evaluation to neutral technical institutions would effectively eliminate the conflict of interest that arises when companies act as both player and referee, boosting the credibility of evaluation findings. Several promising efforts are already underway: academically, AI labs at MIT, Stanford, and other universities provide non-commercial red-teaming; institutionally, Metr (formerly ARC Evals) focuses on frontier model autonomy and dangerous capability assessments, having participated in pre-release testing for GPT-4 and Claude 2; and on the policy side, the EU AI Act requires "general-purpose AI" models to undergo certified evaluations, backed by a dedicated scientific panel. These initiatives offer valuable reference points for building a more systematic third-party evaluation ecosystem.
Strengthen regulatory capacity. Governments need sustained investment to attract top AI talent into the regulatory system. Only then will agencies be genuinely equipped to understand and effectively evaluate complex frontier models.
Finding the Balance Between Innovation and Safety
Transparent, rigorous governance does not mean stifling innovation. A sound AI safety governance framework should provide meaningful security guarantees while avoiding unnecessary obstacles to technological development. The key is striking the right balance — neither turning a blind eye nor allowing excessive caution to impede legitimate technological progress.
Conclusion
"What conversations between governments and AI companies actually look like remains unclear" — that statement alone should serve as a warning signal. In an era of rapidly advancing AI capabilities, a vague understanding of safety evaluation processes is simply not good enough.
Safety decisions about frontier AI models affect everyone, and they should not be made quietly inside a black box. Driving transparency in the evaluation process and establishing accountable governance mechanisms are necessary foundations for the healthy development of AI technology. Only when the public genuinely understands how governments decide whether a model is safe to release can society build the kind of reasonable, sustainable trust this technology requires.
Key Takeaways
Related articles

Disaster and Glory of the Apollo Program: The History We Must Revisit Before Returning to the Moon
From the fatal Apollo 1 fire to Apollo 8's daring lunar orbit to Apollo 11's successful landing—revisiting the disasters, fears, and compromises of the Apollo program and their lessons for today's return to the Moon.

Netflix Trust Exercise Turns Into Firing Trap: Where Are the Boundaries of Corporate Trust?
A Netflix employee was fired after sharing private info in a trust exercise. We analyze the risks of corporate trust exercises and how employees can protect themselves.

AMD CDNA5 Architecture Deep Dive: Technical Evolution and the AI Computing Competition Landscape
Deep analysis of AMD's CDNA5 architecture covering Chiplet packaging upgrades, HBM memory evolution, and low-precision compute optimization, examining how AMD challenges NVIDIA's AI chip dominance.