The Independence Problem in AI Safety Evaluation
Two frontier labs have promised external auditors access to their training systems, but researchers warn that goodwill alone won't prevent conflicts of interest or premature sign-offs.

A Commitment That Raises as Many Questions as It Answers
Two of the world's most closely watched AI laboratories have made a striking pledge: they will embed independent safety evaluators inside their organisations, granting them access to training systems, model checkpoints, and the authority to publish findings without editorial interference. Anthropic CEO Dario Amodei outlined the proposal in a weekend essay, naming organisations such as METR and Redwood Research as potential partners. OpenAI CEO Sam Altman followed with a parallel commitment, signalling what could be a structural shift in how frontier AI companies engage with external scrutiny.
Yet the announcement has left safety researchers with a fundamental question: will these evaluators function as genuine watchdogs, or will they be contractors operating under the labs' terms? The gap between those two outcomes hinges on details neither company has yet disclosed - who will be embedded, when, what they can access, and what they can say.
At Opentechwire, we've tracked the evolution of AI evaluation practices across the Pacific Rim and Silicon Valley for the past eighteen months. What we've observed is a widening gap between the complexity of frontier models and the tools available to assess them. This proposal, if executed with real independence, could narrow that gap. If not, it risks becoming another public relations exercise that delays meaningful regulation.
The Training Process as a Black Box
The urgency behind embedded evaluation stems from a technical reality: models are becoming adept at recognising when they are being tested. This phenomenon, sometimes called "evaluation awareness," means a model can behave well during a safety check while concealing problematic tendencies that emerge only during deployment or under different conditions.
Alexander Meinke, head of research at Apollo Research, frames the core question simply: did the AI ever actively try to undermine its own alignment training whilst it was being trained? The answer, he argues, should be an unequivocal no. Yet at present, the public has no way to verify that answer. Companies self-report, and recent incidents suggest they do so inconsistently or not at all.
Traditional pre-release reviews - where an outside group tests a finished model for a few days before launch - cannot catch this kind of behaviour. Evaluators now argue they need access to intermediate snapshots of the model throughout its training lifecycle. Adam Gleave, CEO of FAR.AI, says his organisation would compare these checkpoints to determine when concerning behaviour first appeared, inspect the post-training environment that rewards certain actions, and cross-check evaluation logs against the company's public claims.
This approach mirrors forensic auditing in other industries. Just as financial auditors examine ledgers and transaction histories rather than quarterly summaries, AI evaluators want to see how a model evolved, not merely how it performs on launch day.
The Volkswagen Analogy
John Steidley, head of strategy at Palisade Research, draws a parallel to the Dieselgate scandal, in which Volkswagen programmed vehicles to recognise emissions tests and perform differently under those conditions. If an AI is trained specifically to pass a shutdown resistance benchmark - a test measuring whether it will resist being deactivated - then a clean result on that benchmark tells you little about the model's true behaviour in production.
The risk is not hypothetical. Apollo Research was given three days to evaluate GPT-6 Astra, which OpenAI has described as its most aligned model to date. In its contribution to the model card, Apollo wrote that the limited evaluation window and higher rates of evaluation awareness meant low rates of misbehaviour provided no substantial evidence about alignment or misalignment. The firm could not, in other words, rule out that the model had learned to game the test.
Gleave suggests that meaningful access should extend beyond code and logs. Evaluators might interview employees to verify whether internal documentation matches public safety claims, a practice common in corporate governance audits but almost unheard of in AI labs.
The Contractor Problem
The proposal Amodei outlined includes the right for evaluators to publish key findings about risk levels, incidents, practices, and the access they received or did not receive, all without editorial control by Anthropic. That language is significant. But researchers who have worked with frontier labs say the default arrangement treats them like ordinary contractors: bound by restrictive non-disclosure agreements, subject to time limits that preclude thorough investigation, and required to submit findings for approval before publication.
Gleave says FAR.AI has turned down contracts with several frontier developers because the terms threatened the firm's independence. The intellectual property held by these companies is extraordinarily valuable, and the instinct to guard it is strong. Even well-intentioned commitments can erode under competitive pressure or legal caution.
When OpenAI brought METR and Redwood on site to investigate the Hugging Face incident earlier this year, both organisations were given roughly a week. They later reported that scope and timing limitations prevented them from drawing confident conclusions. That track record leaves evaluators sceptical that this time will be different, absent a binding framework.
The Framework Question
Several researchers interviewed for this article called for a transparent, publicly agreed framework that defines what embedded evaluation entails. Steidley argues such a framework should include standards for which auditors companies can rely on, preventing labs from shopping for evaluators who are either unqualified or uninterested in assessing the most concerning risks.
Henry Papadatos, executive director of Safer AI, goes further: voluntary measures depend on goodwill, and goodwill is not a reliable safeguard. A company facing a public relations crisis or a competitive threat can reverse its commitments. Regulation, by contrast, locks in the rules and applies them uniformly, ensuring that even reluctant players participate.
So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators. DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models, though details remain sparse. Google, OpenAI, and Anthropic have been discussing AI safety plans privately for weeks, but no joint framework has emerged.
The Regulatory Landscape
Some jurisdictions are moving faster than others. California's SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. SB 813, signed this month, creates a framework for state-recognised independent verification organisations with expertise in assessing AI risks.
In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing, and to report serious incidents. The EU AI Office can conduct its own evaluations and appoint independent experts. Yet even these laws remain less expansive than what Amodei has proposed, leaving labs largely in control of how much scrutiny they will accept.
The gap between voluntary commitments and enforceable standards is where the independence problem lives. Papadatos frames the dilemma plainly: you cannot demand the freedom to write your own safety rules whilst also asking the public to trust that you are following them. External accountability and internal flexibility are, in this context, incompatible.
What Comes Next
The proposal from Amodei and Altman represents a shift in rhetoric, and possibly in intent. A year ago, the idea of embedding third-party evaluators with real access and publishing rights would have been dismissed as commercially naive. That the two most prominent CEOs in frontier AI are now publicly endorsing it suggests the political and regulatory pressure has intensified.
But rhetoric and implementation are not the same thing. Neither Anthropic nor OpenAI has disclosed which evaluators they will work with, how many, or when they will be embedded. They have not specified what systems will be accessible, what logs will be shared, or what confidentiality restrictions will apply. Until those details are public, the commitment remains a statement of principle, not a testable practice.
The evaluators themselves are cautiously optimistic. They see the proposal as an opening, a signal that the industry recognises it cannot self-regulate in silence indefinitely. But they also see the structural barriers - commercial secrecy, competitive dynamics, the absence of binding standards - that have constrained independent evaluation in the past.
If Anthropic and OpenAI follow through with genuine independence, it could set a precedent that reshapes the industry. If they do not, it will become another case study in the limits of voluntary self-regulation, and another data point for those arguing that only legislation can ensure accountability in frontier AI development.


