AI agents · OpenClaw · self-hosting · automation

Quick Answer

Claude Leads 26% of Anthropic's AI R&D: What It Means

Published:

The headline number

On September 17–18, 2026 Anthropic’s policy institute published Measurements for understanding the pace of AI development inside frontier labs. Its most-quoted line: as of August 2026, Claude “leads” 26% of Anthropic’s AI R&D work. Two companion numbers matter as much:

  • The share of work at or above “AI collaborates” is above 90%.
  • Claude is not operating fully autonomously for any measured subset of AI R&D work.

So the honest one-sentence summary is: most of Anthropic’s model-building now involves Claude doing large chunks of the work; in about a quarter of it Claude runs the task and a human supervises; in none of it is the human absent.

The scale behind the number

Anthropic used the Automation Level (AL) scale developed by Epoch AI, which runs from AL0 to AL5:

LevelLabelWhat it means
AL0No AIHumans do the task
AL1–AL2AI assistsAI helps with parts of the task
AL3AI collaboratesAI does large chunks under close human direction
AL4AI leadsAI completes most of the task end-to-end from a high-level prompt; human supervises
AL5AI autonomousNo human in the loop

Anthropic built a prototype R&D Automation Index by cataloguing every kind of AI R&D task at the company, rating each one’s current AL, and aggregating. The 26% is the share at AL4; the “>90%” is the share at AL3 or above; the “none” is AL5.

How fast it moved

The trajectory is the alarming part, and it is why the story led the AP, the Los Angeles Times, Business Insider and the Washington Post. Reporting on the release puts the AL4 share at effectively zero in February 2026 and below 1% in March 2026 — six months before it hit 26%. Anthropic says explicitly that it publishes the metric to track how close the world is to recursive self-improvement, defined as a model fully autonomously building its successor, and that “models accelerating their own development could make it more challenging for humans to understand or control these systems.”

What the measurement leaves out

Anthropic is unusually candid about the limits:

  1. No common methodology. Nobody else measures this, so there is nothing to compare against.
  2. The judge is the subject. Anthropic uses its own models to rate how automated its own work is, so the rater could share the blind spots of the thing it rates.
  3. It is a prototype. The catalogue and ratings are internal and unverified.

Its proposed fix is embedding independent third-party evaluators from multiple organisations inside Anthropic, with access comparable to internal risk-assessment teams, to verify safety practices, report incidents and monitor these metrics.

The other two measurements

The 26% figure is one of three proposed yardsticks:

  • Oversight of AI agents — how well Anthropic can see and intervene in actions agents take on its own systems, as employees increasingly delegate to Claude.
  • Compute allocation — how the resources that power model development are split between training, experiments, safety testing and inference.

Together with capability evaluations (published through the Responsible Scaling Policy risk reports) and the transparency obligations in Anthropic’s Advanced AI Framework policy proposal, they are pitched as a starting point for monitoring the pace of AI development from outside the labs.

Why now

The post landed in a loaded week. CEO Dario Amodei published a call to slow development of the most powerful systems (“we must pace the frontier”), which Sam Altman and Elon Musk publicly backed and President Trump dismissed. Three days earlier Zuckerberg, Musk and Huang had persuaded the White House to drop an industry-funded AI regulator. And OpenAI, in its own September disclosures, described a milestone of models acting as automated research interns. The 26% number is Anthropic’s evidence that the acceleration is real and its argument that measurement, not trust, should govern what happens next.

How to read it as a builder

  • For engineering teams: AL4 at 26% inside the lab that makes the model is a preview of where agentic coding is heading for everyone else — the coordinator-plus-workers pattern in Claude Code Projects is the productised version.
  • For policy watchers: the numbers are voluntary, self-rated and one lab deep. Their value is the precedent: a frontier lab publishing a repeatable pace metric and inviting others to match it.
  • For sceptics: “leads” is not “autonomous.” The claim is narrower than the headlines, and Anthropic says so.

Sources