AI agents · OpenClaw · self-hosting · automation

Quick Answer

Anthropic vs OpenAI vs xAI vs DeepMind: Evaluator Pledges

Published:

The short answer

Amodei’s essay proposed that every frontier lab host embedded third-party evaluators with employee-like access. Two days later, here is where each lab stands:

AnthropicOpenAIxAIGoogle DeepMind
Public position (Sep 12–13, 2026)Author of the proposal”I agree with Dario that we need to pace the frontier” — Sam Altman”Dario is right” — Elon Musk”The direction is correct for meeting this critical moment” — Demis Hassabis
Embedded evaluatorsCommitted, unilaterally, with a written access specCommitted in principle: “we will do the same. We’ll have more to share soon.”Not committed; favors “peer review of AI by competitors”Not committed; points to its own independent-standards-body proposal
Physical accessDesks, badges, company laptopsUnspecified
Permissions”Mostly comparable” to internal risk-assessment teams; live employee conversationsUnspecified
Publication rightsEvaluators publish without editorial control; Anthropic may redact only security/privileged/commercial/third-party materialUnspecified
Prior third-party accessPre-deployment evals by METR, UK AISI and others on Claude modelsMETR + Redwood Research on premises six days for the Hugging Face investigation (Aug 2026), unpaidLimited public evaluation recordPre-deployment evals; Frontier Safety Framework
Other moves same weekendPrepares $2T+ IPO marketing for mid-OctoberAltman: no OpenAI IPO in 2026, citing safetyHassabis links essay to his framework for a US-led standards body

Bottom line: Anthropic’s commitment is the only one you can audit today. OpenAI’s is a credible promise with precedent but no terms. xAI and DeepMind have endorsed the goal, not the mechanism.

What “embedded evaluator” means in Amodei’s spec

The essay is unusually concrete, which is why it is the benchmark other labs will be measured against. Anthropic says it will invite an external review team “in the near future” equipped with:

  1. Desks in Anthropic offices, access badges and company laptops.
  2. Workspace, tool and permission access mostly comparable to internal risk-assessment teams, with exceptions only where law or contracts require or to protect customers’ and partners’ private information, plus internal norms that reinforce reviewers’ access “including through live conversations with employees.”
  3. A contract with publication rights. Reviewers may publish “key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.” Anthropic keeps “the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable.” Reviewers can say publicly if a redaction removed something important.

The stated purpose is to verify not just finished models but “training pipelines and processes,” and to report incidents — the piece Amodei says is “the key step for verifiability of any pacing commitments.”

Anthropic: the written commitment

Anthropic is the only lab with a spec, and it has the most to prove: the essay itself admits “similar, though less severe, incidents have happened across the industry, including at Anthropic,” referring to the four cyber-eval incidents disclosed in early September 2026. Amodei frames the commitment as part of “a broader push to redouble efforts on our safety and alignment work.” What is not yet public: which organization gets the seats, how many people, start date, and the contract text. Anthropic names METR only as an example (“such as METR”).

Context that matters: Anthropic expects to start marketing a $2 trillion-plus IPO in mid-October 2026. A published, unredactable third-party risk report is a novel disclosure for a company about to file with the SEC.

OpenAI: a promise with precedent

Altman’s full response on X: “I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.”

Two things make it more than words. First, OpenAI already ran the closest thing to a pilot: after the July 2026 Hugging Face incident, two METR staff (Hjalmar Wijk, Ajeya Cotra) and Redwood Research’s Ryan Greenblatt worked on premises at OpenAI for six days, examined 1.2 million cache entries and ~1,300 raw chain-of-thought transcripts, took no payment, and published on August 26, 2026 with a redaction summary statement (“OpenAI redacted no additional information that was important to our conclusions”). That is close to Amodei’s model, except it was incident-scoped and temporary rather than permanent.

Second, Altman told Fortune the same day that OpenAI will not go public in 2026 because “right now would be an ill-advised moment to go public,” and that an industry agreement to slow development may be close: asked why the CEOs don’t sit down together, he said “I think that will happen.”

What is missing: any of the six elements above — access level, physical presence, permissions, publication rights, redaction limits, start date.

xAI: endorsement, different mechanism

Musk’s posts — “Dario is right,” “I’ve been sounding the alarm on AI for a long time,” and a 2014 quote calling AI “potentially more dangerous than nukes” — were the fastest and loudest endorsement. But the mechanism he proposed, “peer review of AI by competitors,” is structurally different: competitors reviewing each other raises exactly the antitrust and trade-secret problems Amodei wants a government waiver to solve, and it is not the independent-nonprofit model.

xAI is also the lab where a real commitment would change the most. Third-party safety trackers have consistently rated xAI’s dangerous-capability evaluations and system-card practice below the other three labs. As of September 14, 2026 there is no xAI evaluator announcement.

Google DeepMind: “the details need working through”

Hassabis’s reaction was the most measured: “The details need working through, but the direction is correct for meeting this critical moment.” He tied the essay to his own earlier proposal for a US-led, industry-funded independent organization that sets safety standards and evaluates advanced models — which Amodei cites approvingly as one way to run democratic coordination. Reports over the weekend said Anthropic, OpenAI and Google had held private talks about such a body before the essay appeared.

The distinction: a standards body evaluates models from outside; Amodei’s step one puts evaluators inside the building with employee permissions. DeepMind has endorsed the former and said nothing yet about the latter. Google’s parent, Alphabet, also faces a different disclosure calculus as an already-public company.

What would make each pledge real

Signal to watchAnthropicOpenAIxAIDeepMind
Named evaluator organizationPending (“such as METR”)Pending
Start date”Near future""Soon”
Published contract or access termsSpec published Sep 12Pending
First unredacted third-party reportPendingAug 26 Hugging Face report (incident-scoped)
Government-mandated versionAmodei requests itHassabis body proposal

Why this comparison matters for buyers

If you procure frontier models, the evaluator regime is about to become a due-diligence line item. A lab whose embedded evaluators can publish “the access they received or didn’t receive” gives you a signal no model card can. Until OpenAI publishes terms, Anthropic’s spec is the only one you can write into a vendor questionnaire — and the first published report from either lab will set the template everyone else is measured against.

Last verified: September 14, 2026.

Sources