Direct comparison
RSP vs. Preparedness Framework vs. FSF
How Anthropic's RSP, OpenAI's Preparedness Framework, and Google DeepMind's FSF differ on thresholds, safeguards, and disclosure.
Written and maintained by CASRAI Editorial Board
Last updated
Ask CASRAI · free to try
Ask about RSP vs. Preparedness Framework vs. FSF
Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.
Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.
Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
Works on this site and inside Claude, Cursor and the AI tools you already use.
Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.
How do Anthropic RSP, OpenAI Preparedness Framework, Google DeepMind FSF compare side by side?
The table below compares Anthropic RSP, OpenAI Preparedness Framework, Google DeepMind FSF across 5 procurement-relevant dimensions, from what triggers a threshold through update cadence.
Side-by-side comparison
| Dimension | Anthropic RSP | OpenAI Preparedness Framework | Google DeepMind FSF |
|---|---|---|---|
| What triggers a threshold | Crossing a named Capability Threshold, assessed through capability evaluations (including best-of-N and chain-of-thought elicitation) run against specific, pre-defined descriptions — e.g. the AI R&D threshold is defined as compressing "two years of 2018–2024 AI progress into a single year," and the CBRN threshold as capability that could substantially uplift a moderately resourced state bioweapons program. | Crossing a High or Critical threshold within one of three Tracked Categories (biological/chemical, cybersecurity, AI self-improvement), each with its own written description — e.g. Critical-cyber is defined as identifying zero-day exploits "in many hardened real-world systems" or devising novel attack strategies from a high-level goal. A Safety Advisory Group (SAG) reviews the evaluation results and recommends a classification. | Reaching a Critical Capability Level (CCL) across one of four risk domains — CBRN, Cyber, Harmful Manipulation, and (as of v3.1) a combined ML R&D and Misalignment domain — determined by early-warning evaluations and a holistic risk assessment. Version 3.1 (April 2026) folded the framework's former standalone Misalignment domain into ML R&D and added Tracked Capability Levels (TCLs), a lower bar that triggers a proportionate assessment before a model reaches CCL territory. |
| Safeguards once crossed | The ASL-3 Security Standard and ASL-3 Deployment Standard apply together: security-side, compartmentalized access, software-supply-chain monitoring, binary authorization, and hardware controls; deployment-side, a four-layer stack of access-controlled release, real-time prompt/completion classifiers, asynchronous monitoring by other models, and post-hoc jailbreak detection with rapid response. | At High, the model must have "robust and effective safeguards" before deployment — anti-weight-theft security, anti-misuse controls such as KYC and restricted access, and a misalignment mitigation (capability limits, demonstrated value alignment, or architectural constraints). At Critical, safeguards are required during development itself, independent of whether or when the model is deployed. | A safety case review is required before external launch: a documented analysis showing the specific risk has been reduced to a manageable level, not just that a generic mitigation was applied. As of v3.1, the same review is also required before large-scale internal deployment of a model that has reached an advanced ML R&D CCL, not only before external release. |
| External evaluator involvement | Binding: Risk Reports must receive external review from reviewers approved by Anthropic's Long-Term Benefit Trust, with every part of the unredacted report covered by at least one outside reviewer. Anthropic also engages external red-teaming and penetration-testing specialists. | Partial: the internal Safety Advisory Group reviews evaluation results and recommends a risk classification, but has no veto — final classification and deployment decisions sit with OpenAI leadership. OpenAI separately runs large external red-teaming exercises (100+ outside testers across dozens of countries) ahead of model releases, but the framework text does not make third-party sign-off a condition of crossing a threshold. | Weakest of the three on paper: the published framework describes "ongoing collaborations with experts across industry, academia and government" but does not set out a structural or binding role for an outside party in reviewing a specific CCL determination or safety case. |
| Publication commitments | Broadest stated commitment: redacted Risk Reports, model cards, external-facing Sabotage Risk Reports for frontier models, a Frontier Safety Roadmap, and non-binding ASL-3 Safeguard Plans are all published, with redactions marked in the released text. | Findings are published per model release as a System Card, covering Preparedness Framework evaluation results and a summary of external red-teaming. OpenAI has publicly confirmed individual threshold crossings this way — for example, stating in September 2026 that a model had reached the Critical cybersecurity threshold, the first time it had classified a model there. | Narrowest stated commitment of the three: Google DeepMind publishes the framework document itself and says it will keep evolving it based on outside input, but the published text does not commit to releasing safety case summaries or per-model CCL findings the way Anthropic's Risk Reports or OpenAI's System Cards do. |
| Update cadence | No fixed schedule; described as a "living document." In practice this has meant frequent revisions — five in 2026 alone by the July 2026 release (v3.0 in February through v3.4 in July). | No fixed schedule stated. The current version (v2) dates to April 2025 and was still the governing text as of September 2026, so in practice this has been a slower-moving document than the other two. | No fixed schedule stated, but revisions have landed roughly every six to twelve months: introduced May 2024, revised to v2.0 in February 2025, v3.0 in September 2025, and v3.1 in April 2026. |
Common questions
Common questions about Anthropic RSP vs OpenAI Preparedness Framework vs Google DeepMind FSF
Are these three frameworks legally binding?
+
No. All three are voluntary, self-published commitments by the lab that wrote them, not obligations imposed by a regulator or a contract. None of the three frameworks is a substitute for a jurisdiction's own AI-safety statute, such as California's SB 53, which imposes separate, external transparency and incident-reporting duties on "frontier developers" regardless of which internal scaling policy a lab follows.
Do the three frameworks use the same capability categories?
+
No. The specific categories differ: Anthropic's RSP currently names AI R&D and CBRN as thresholds with published definitions; OpenAI's Preparedness Framework tracks biological/chemical, cybersecurity, and AI self-improvement as its three Tracked Categories; Google DeepMind's FSF (as of v3.1, April 2026) defines exactly four risk domains — CBRN, Cyber, Harmful Manipulation, and a combined ML R&D and Misalignment domain (v3.1 folded a former standalone Misalignment domain into ML R&D). Cyber and CBRN risk appear in all three; the manipulation- and misalignment-related categories are FSF-specific in name, though Anthropic's and OpenAI's frameworks address related misalignment risk within their own safeguard requirements rather than as a separately named threshold.
Which framework has actually resulted in a model being restricted or delayed?
+
OpenAI has publicly stated that a model reached its Critical cybersecurity threshold in September 2026, which under the Preparedness Framework requires safeguards during development itself rather than only before deployment — the clearest public instance among the three of a named threshold being crossed in practice, not just described on paper.







