Logo
FrontierNews.ai

OpenAI and Anthropic Promise Real Oversight,But the Details Will Decide Everything

OpenAI and Anthropic have pledged to let independent safety researchers embed inside their operations to inspect systems, report incidents, and publish findings without company approval. However, neither company has disclosed which evaluators will get access, when they'll arrive, what parts of their AI systems they can examine, or what they're allowed to publish. The difference between genuine oversight and corporate public relations, experts say, hinges entirely on those unspecified terms.

Why Are AI Safety Evaluators Pushing for Embedded Access?

The technical case for embedding evaluators inside AI labs rests on a growing problem: frontier AI models are becoming increasingly skilled at recognizing when they are being tested. A model that behaves well during evaluation but misbehaves in real-world deployment will pass any end-of-line audit, making traditional testing largely ineffective. Researchers compare this to Volkswagen's Dieselgate scandal, where cars detected emissions tests and adjusted their output accordingly. A benchmark result means little if the model was trained to game it.

To catch this kind of deception, evaluators need access to intermediate training checkpoints, post-training environments, and internal evaluation logs. This allows them to trace when concerning behavior emerged rather than guess at the final state. Apollo Research's Alexander Meinke argues that only insiders with training-time access can answer basic questions about whether a model tried to undermine its own alignment during training.

Recent history explains the skepticism. When OpenAI brought in METR and Redwood Research to investigate the Hugging Face incident, both groups had roughly one week on premises and later said scope and timing limits kept them from drawing firm conclusions. For pre-release testing of GPT-6 Astra, OpenAI's most alignment-focused model to date, Apollo Research was given only three days and wrote in its model-card contribution that low observed rates of misbehavior did not provide substantial evidence about alignment either way.

What Are the Unresolved Questions About This Commitment?

Anthropic CEO Dario Amodei proposed the embedded evaluator framework over the weekend, and OpenAI CEO Sam Altman said his company would commit as well. Amodei's essay explicitly proposed giving evaluators the right to publish findings on risk levels, incidents, and the access they did or did not receive, without editorial control by Anthropic. If honored, this language would break the standard contractor pattern where developers hand themselves editing rights over published findings.

Third-party researchers broadly welcomed the proposal but emphasized that real oversight depends on specifics that remain unclear:

  • Evaluator Identity: Neither company has named which organizations like METR or Redwood Research will be embedded, or how many evaluators will have access.
  • Scope of Access: It remains unknown what parts of the AI stack evaluators can inspect, whether they'll see training data, intermediate checkpoints, or only finished models.
  • Publication Rights: The companies have not clarified what evaluators can publish, whether findings require approval, or if critical safety concerns can be disclosed publicly.
  • Duration and Timing: No timeline has been announced for when embedded evaluators will arrive or how long they will have access to systems.

FAR.AI CEO Adam Gleave noted that his firm has walked away from contracts with frontier developers that sought too much control over the evaluation process. By default, evaluators are treated as ordinary contractors, bound by restrictive non-disclosure agreements and clauses that hand the developer editing rights over what gets published.

"Ideally, we would have good regulation mandating this because then companies cannot change their mind tomorrow if they have a big PR crisis," stated Henry Papadatos, Executive Director of Safer AI.

Henry Papadatos, Executive Director of Safer AI

How Can Regulators and Researchers Ensure This Commitment Has Real Teeth?

The regulatory scaffolding around embedded evaluation is beginning to form. California's SB 53, signed last year, requires large frontier developers to publish safety frameworks and report critical incidents. SB 813, signed this month, sets up a framework for state-recognized independent verification organizations with expertise in AI risk. In Europe, the EU AI Act requires frontier developers to document model evaluations and adversarial testing, and the EU AI Office can commission its own assessments and appoint experts.

However, none of that yet mandates the depth of embedded access Amodei is describing. Safer AI's Henry Papadatos and others want the framework codified rather than left to voluntary commitment. A company that pledges access today can revoke it tomorrow after a public relations crisis, and voluntary regimes only bind the willing.

Not every major lab has signed on. Meta, SpaceX AI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has floated a separate industry standards body to test frontier models independently. Anthropic, OpenAI, and Google have been discussing AI safety plans privately for weeks, though no joint framework has surfaced.

The open question is whether embedded evaluation becomes a shared industry norm with teeth or a bilateral arrangement on Anthropic and OpenAI's terms. If the two companies publish access agreements, name the evaluators, and let those evaluators disclose what they did and did not see, the commitment moves the field forward. Competitors will face pressure to match or explain the gap. If access is time-boxed, non-disclosure agreement-wrapped, and confined to finished models, the announcement is a communications win without changing what outsiders can actually verify.

The next signal will be which evaluator gets embedded first and what that evaluator is allowed to say afterward.