AI Giants Plan to Embed Independent Safety Evaluators — But Can 'Independence' Be Guaranteed?

Anthropic and OpenAI's plan to embed independent safety evaluators wins cautious praise, but independence remains unproven.
Anthropic and OpenAI are pursuing a rare move: allowing independent safety evaluators inside their labs with direct access to frontier model development. Researchers respond with cautious optimism, raising concerns on three fronts — transparency (access doesn't guarantee public accountability), independence (structural conflicts of interest may compromise evaluators' judgment), and long-term governance (voluntary arrangements lack binding force and can be revoked). Researchers broadly see embedded evaluators as a transitional step, but argue that reliable AI safety governance ultimately requires institutionalized external oversight, not corporate self-regulation.
An Unprecedented Opening of AI Labs
Anthropic and OpenAI are pushing a move that is remarkably rare in the industry — embedding independent safety evaluators directly inside their AI laboratories. This would give external researchers direct access to frontier model development processes, a level of transparency that was nearly unimaginable until now.
For the research community that has long called for stronger AI oversight, this is undoubtedly a positive signal. The capabilities of frontier large language models are evolving far faster than outside observers can track, while training data, evaluation methodologies, and risk mitigation measures inside these labs have typically been treated as core trade secrets. Allowing independent evaluators in at least formally breaks down that wall of secrecy.

Researchers Welcome the Move — With Caution
The research community's reaction can be summed up as "welcome, but trust is not given freely." Access that was previously unimaginable is certainly worth acknowledging — but access alone does not equal effective oversight. Several researchers have noted that truly meaningful safety evaluation requires three conditions to be met simultaneously: transparency, independence, and ultimately, institutionalized regulation.
Transparency: Access Doesn't Mean Being Fully Informed
Being embedded inside a lab is only the first step. If evaluators only see curated information, or if evaluation findings cannot be disclosed to the public, the practical value of such "access" is severely diminished. Transparency demands not just that evaluators can see the models, but that the evaluation process and conclusions can be subjected to external scrutiny.
Independence: Who Pays the Piper, Calls the Tune?
The most fundamental skepticism centers on the word "independent." When evaluators are invited by the very lab being evaluated — and may even be resourced by it — could their judgment be influenced by implicit conflicts of interest? This is a structural problem. True independence typically requires that an evaluating body have full control over its personnel, funding, and publication of findings, free from the influence of the party being evaluated. Whether the current embedded arrangement can achieve this remains an open question.
This dilemma is well known in regulatory circles, where academics call it "regulatory capture" — the phenomenon whereby the regulated party gradually influences or even co-opts the oversight body meant to hold it accountable, through superior resources, information, or relationships. In the context of AI safety evaluation, regulatory capture could occur in more subtle ways: evaluators immersed in lab culture for extended periods may unconsciously absorb the lab's own risk perception framework, or they may hedge their conclusions out of a practical desire to preserve their access. Historically, credit rating agencies in finance and pay-to-review systems in drug approval have both drawn heavy criticism for similar structural flaws. Avoiding this risk typically requires careful design across multiple dimensions — evaluator selection processes, diversified funding sources, and independent channels for publishing findings.
From Voluntary Commitments to Mandatory Oversight
Researchers broadly agree that while companies voluntarily bringing in evaluators is progress, reliable long-term AI safety governance cannot do without external regulatory intervention. Voluntary arrangements carry an inherent fragility — they can be tightened or cancelled at any time, and they lack binding enforcement.
Only when independent evaluation is embedded in a legal or mandatory industry framework will evaluators' access rights, independent status, and disclosure obligations be truly guaranteed. In other words, embedded evaluators can serve as a transitional step toward institutionalized oversight, but should not be treated as the destination.
A handful of jurisdictions around the world have already begun exploring mandatory AI evaluation frameworks. The EU's Artificial Intelligence Act (EU AI Act) requires high-risk AI systems to undergo compliance assessments before market deployment and establishes transparency and third-party audit obligations for "general-purpose AI models" (GPAI). The United States, by contrast, has relied primarily on executive orders and voluntary commitments, and has yet to develop a legally binding, unified evaluation regime. Against this backdrop, the voluntary initiatives from Anthropic and OpenAI carry short-term demonstrative value — but they also highlight a glaring institutional vacuum: in the absence of a mandatory framework, companies can choose to open up, but they can just as easily choose to close ranks, with external parties having almost no legal means of intervention.
An Experiment in Trust
At its core, this move by Anthropic and OpenAI is an attempt to trade "internal openness" for external trust in their safety commitments. It reflects a proactive posture from frontier AI companies responding to pressure for public accountability, while also exposing the fundamental dilemma of current AI governance: in the absence of a mature regulatory framework, the credibility of safety evaluations depends heavily on corporate self-restraint.
The success or failure of this experiment will depend on the answers to several critical questions — Can evaluators achieve genuine independence? Can evaluation findings be made public? And will the industry move toward mandatory regulation? Until these questions are clearly answered, "independent safety evaluators" looks more like a promising but yet-to-be-proven beginning.
Related articles

rag-eval: A Zero-Dependency, No-API-Key RAG Evaluation Tool
rag-eval is a zero-dependency, framework-agnostic open-source RAG pipeline evaluation tool. It supports free local lexical and retrieval metrics with no API keys required, and offers optional LLM Judge for semantic validation. Compatible with Haystack, LangChain, and LlamaIndex.

Vercel AI SDK Releases workflow-harness 1.0.115 Patch Update
Vercel AI SDK releases @ai-sdk/workflow-harness 1.0.115 patch update, syncing the @ai-sdk/harness dependency. Learn about the update, release mechanism, and what it means for developers.

GLM 5.3 Now Available on Serverless Training API — No Sales Process Required
GLM 5.3 is now available on Serverless Training API alongside Kimi K3 and Qwen 3.8 27b. No sales process needed — start fine-tuning directly via docs or pre-made recipes.