OpenAI vs Anthropic vs Google: AI Safety Rules 2026
The Short Answer
Three voluntary frameworks, broadly convergent on what to measure and divergent on how much they promise. None is externally enforced.
| OpenAI | Anthropic | Google DeepMind | |
|---|---|---|---|
| Framework | Preparedness Framework | Responsible Scaling Policy | Frontier Safety Framework |
| First published | 2023 | Sept 2023 | 2024 |
| Notable version | v2.0 (Apr 2025), extensions signalled Aug 2026 | v3.1 (Apr 2026) | v3.0 (Apr 17, 2026) |
| Threshold language | Capability categories, High / Critical | AI Safety Levels (ASL) | Critical Capability Levels (CCLs) |
| Distinctive element | Public incident disclosure | Weight-security standards by level | Tracked Capability Levels; Harmful Manipulation CCL |
| Core domains | Bio/chem, cyber, AI self-improvement | Bio/chem, cyber, AI R&D | Bio/chem, cyber, manipulation, ML R&D |
| Enforcement | Self | Self | Self |
Verified August 19, 2026. Frameworks are revised frequently — check the vendor page for current version.
What These Documents Actually Are
A frontier safety framework is a public, unilateral commitment: the lab states which dangerous capabilities it will test for, what level of capability triggers additional safeguards, and what those safeguards are. They are not regulation. No government approves them, no auditor signs them, and each lab can revise its own.
That sounds weak, and in an important sense it is. But it is not nothing, for one reason: revisions are visible. When a lab weakens a commitment, external researchers notice and write it up. Reputational cost is the enforcement mechanism.
Independent work exists to compare them. METR maintains an analysis of common elements across frontier safety policies, and SaferAI assessed the frameworks of twelve companies — Amazon, Anthropic, Cohere, Google DeepMind, Magic, Meta, Microsoft, Naver, Nvidia, OpenAI, xAI and others — against 65 weighted criteria across risk identification, analysis, treatment and governance. The consistent finding: broadly similar, misuse-focused approaches with meaningful divergence in specificity and in what triggers a genuine stop.
Where They Converge
All three track roughly the same risk domains, and this convergence is itself informative:
- Biological and chemical uplift
- Cybersecurity capability — the domain that dominated 2026
- AI self-improvement / AI R&D acceleration
All three use a tiered structure: measure capability, cross a threshold, trigger stronger safeguards on both security (protecting weights from theft) and deployment (restricting what the model will do for users).
All three have, at some point, relaxed requirements they had previously set. Anthropic and Google DeepMind each reduced safeguards for some CBRN and cybersecurity capabilities after concluding initial requirements were excessive; OpenAI removed persuasion from the Preparedness Framework entirely, handling it through other policies. Whether you read that as sensible calibration or as commitment erosion is the central disagreement in this field.
Where They Diverge
Google DeepMind — the widest scope. FSF v3.0 (April 17, 2026) added two things its peers lack: a Tracked Capability Levels early-warning layer that flags capabilities approaching a threshold rather than only at it, and a Harmful Manipulation CCL governing models capable of systematically altering human beliefs. Against Anthropic’s RSP v3 revisions and OpenAI’s Preparedness v2 simplification, DeepMind’s is the framework that expanded rather than narrowed.
Anthropic — the most specific on weight security. The RSP’s AI Safety Level structure ties concrete security standards to capability tiers, making “what protections apply at this level” more answerable than in the alternatives. Anthropic has revised the RSP repeatedly, reaching v3.1 in April 2026 — high iteration rate, which cuts both ways.
OpenAI — the most publicly stress-tested. OpenAI’s framework produced a visible, costly action on August 18, 2026: a two-week pause in RL training on deployment-bound models and an ongoing hold on its largest planned frontier run, triggered by preliminary evidence that its Astra model may meet the Critical cybersecurity capability threshold, compounded by the OpenAI-Hugging Face incident. OpenAI also stated that the situation calls for “a broader approach — one that builds on and extends beyond the current Preparedness Framework.”
That last sentence is the most interesting in any of these documents this year. A lab saying its own framework is insufficient is unusual, and it is either commendable candour or an admission that the framework was not calibrated for the capability curve it is now on. Both readings can be true.
What This Means If You Are Buying
Most people reading a framework comparison are not policy researchers; they are deciding which provider to build on. Three practical translations:
1. Frameworks predict availability shocks. A framework that can pause development can also restrict deployment. Anthropic restricted Claude Fable 5 access on cybersecurity grounds in 2026; OpenAI paused Astra. If your product depends on one model with no fallback path, capability-triggered restrictions are a real availability risk. Keep a second provider integrated and tested — this is the strongest practical argument for a model-agnostic gateway.
2. Cyber capability restrictions will hit security tooling hardest. If you build in the security space, expect the most friction, the most access reviews, and the highest chance that a model you rely on becomes conditionally available.
3. The frameworks say nothing about your security. Every one of them governs the lab’s models, not your deployment. The August 2026 incident chain — malicious dataset, code execution, credential harvesting, lateral movement — was stopped by none of these documents. Your sandboxing, credential scoping and logging are the controls that apply to you.
The Honest Assessment
These are voluntary commitments by commercial organisations in a competitive race, revisable by the organisations themselves, verified by no one external. That is a weak form of governance.
It is also, as of August 2026, the only form that has demonstrably delayed a frontier training run. OpenAI incurred real cost against a written commitment and published the reasoning. Judge frameworks by whether they produce costly actions under pressure, not by how comprehensive the document reads. On that test there is now exactly one public data point, and it is worth more than the other three thousand words of comparison above.
Last verified: August 19, 2026.