Skip to main content
v2026.11,858 entries · CC-BY 4.0

Marginal Risk vs. Absolute Risk: The Contested Framing Behind Frontier AI Safety Cases

Anthropic’s RSP rewrote itself around a distinction between a model’s absolute risk and its marginal risk relative to what other developers already field — and built in extra governance friction specifically for when that comparison drives a decision. OpenAI, Google DeepMind, and Meta all reason the same way, but each anchors the comparison to a different baseline. This guide sets out what each framework actually says, in its own words, and where the comparison holds up under scrutiny.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: September 20, 2026. “Marginal risk” means how much danger a company’s own model adds on top of what’s already achievable with other AI systems already on the market; “absolute risk” means the total danger the model poses on its own, measured against no such comparison. Anthropic’s Responsible Scaling Policy (RSP) v3.4, effective July 8, 2026, explicitly argues from both — and treats the choice between them as consequential enough to require extra board-level sign-off. OpenAI, Google DeepMind, and Meta each build a version of the same comparison into their own frameworks, but anchor it to different baselines, and the Frontier Model Forum has named the specific failure mode — “risk creep” — that a purely marginal basis permits.

What the two terms actually mean, in the frameworks’ own words

Anthropic’s RSP v3.4 defines the pair directly, in a footnote to its Risk Reports section: “By ‘marginal risk analysis,’ we mean arguing that the risks imposed by our systems in particular are relatively lower when keeping in mind the risks unavoidably posed by other AI systems.” Absolute risk, by contrast, is what’s left over once a model’s own mitigations are accounted for — the RSP calls it “remaining absolute risk—i.e., the leftover risk after accounting for our mitigations,” assessed for each threat model independently of what any other developer is doing.

The distinction is not academic within the policy itself. Section 3.2 commits Anthropic to “acknowledge when we view certain models as posing significant risks in absolute terms, even if our marginal contribution to overall ecosystem risk may be relatively limited when taking other developers’ AI models into account” — meaning the RSP does not let a favorable marginal comparison silently override an unfavorable absolute one.

Why Anthropic rewrote its own RSP around this distinction

This is a real change, not a restatement of prior practice, and the RSP says so directly: “Our previous RSP committed to implementing mitigations that would reduce our models’ absolute risk levels to acceptable levels, without regard to whether other frontier AI developers would do the same. But from a societal perspective, what matters is the risk to the ecosystem as a whole.” The stated reason is a collective-action problem: “If one AI developer paused development to implement safety measures while others moved forward with training and deploying AI systems without strong mitigations, that could result in a world that is less safe — the developers with the weakest protections would set the pace.”

Because the policy treats marginal-risk reasoning as a real risk to the reasoning itself, not just a convenient argument, it attaches procedural friction to it. Section 3.3 requires that whenever “the absolute level of risk being imposed industry-wide by AI systems such as ours is high, and are justifying our decision to move forward based partly on a marginal risk analysis,” Anthropic must additionally disclose a competitive-landscape analysis, how that comparison factored into the risk assessment, a benefits analysis, and its own advocacy efforts to raise the underlying risk with regulators and peers. Section 3.4.5 goes further: “In the event that marginal risk analysis… plays a major role in a decision to move forward, explicit approval of the Risk Report by the Board and LTBT (rather than just the CEO and RSO) will be required” — the company’s two highest internal governance bodies, rather than its normal sign-off chain.

Three other labs make the same comparison — against three different baselines

Anthropic is not alone in reasoning this way, and NIKOLAI’s crosswalk of the concept (below) finds at least four distinct baselines in active use, none interoperable with the others:

  • Anthropic compares a model’s risk to “the risks unavoidably posed by other AI systems” currently on the market — a moving, competitor-relative baseline.
  • OpenAI’s Preparedness Framework v2 (§4.3) reasons the same way — “we could adjust accordingly the level of safeguards that we require,” but only if OpenAI can “rigorously confirm” a competitor has shipped a comparably capable, less-safeguarded system, and even then only while keeping its own safeguards “at a level more protective than the other AI developer, and share information to validate this claim,” explicitly “in order to avoid a race to the bottom on safety.” Several of its own capability thresholds are separately defined against a fixed 2021 baseline: whether a model provides “meaningful counterfactual assistance (relative to unlimited access to baseline of tools available in 2021).”
  • Google DeepMind’s Frontier Safety Framework v3.1 uses a competitor-relative comparison for its risk-acceptance test — “if other models are similarly capable and have few mitigations, then the marginal risk added by our external deployment is likely low” — but defines its capability thresholds themselves against a fixed baseline: “relative to a baseline without generative AI” (footnote 8).
  • Meta’s Advanced AI Scaling Framework v2 uses a “net new” test set against general-purpose AI’s existence at all: a risk counts only if “the outcome cannot currently be realized as described… with existing tools and resources but without access to general-purpose AI.”

So a frontier developer’s marginal-risk claim can mean “safer than my closest competitor,” “safer than the internet circa 2021,” or “not something generative AI made newly possible at all” — three genuinely different comparisons, each defensible on its own terms, none directly commensurable with the others.

The warning built into the industry’s own coordinating body

The Frontier Model Forum — the industry body Anthropic, OpenAI, Google DeepMind, Microsoft, Meta, Amazon, and others fund jointly — names the systemic risk of exactly this framing in its Risk Taxonomy and Thresholds report (§3.2, Establishing Risk Baselines). It defines “Marginal Risk Assessments” as developers “consider[ing] the additional risk their model adds beyond what’s already possible in the ecosystem,” then states the failure mode directly: such assessments “may facilitate gradual risk escalation through ‘risk creep’ — where multiple models, each introducing only small marginal increases which do not cross individual ‘net new’ thresholds, may collectively create significant risk growth.” The alternative — a static, historical threshold fixed at one point in time — “provide[s] consistency and clear benchmarks” but “could handicap development if other actors proceed with less stringent safeguards.” The Forum names the trade-off; it does not resolve it, and neither, on their own telling, do the labs that fund it.

Where Narayanan and Kapoor actually stand — and where they don’t

CASRAI has already covered Princeton researchers Arvind Narayanan and Sayash Kapoor’s “AI Safety Is Not a Model Property” (March 12, 2024) on this site, and it is tempting to read them as the “absolute risk” side of this debate. That is not quite what the primary source says, and it is worth being precise rather than convenient. In that same essay, Narayanan and Kapoor actually endorse rigorous marginal-risk analysis as the more defensible approach to one specific question — whether to release a model’s weights openly — citing a Stanford framework that “enables assessing the marginal risk of releasing a model… compared to the risk from existing models,” and reporting that this method found open models’ marginal cybersecurity risk low while their marginal risk for generating non-consensual intimate imagery was “substantial.” Their objection is not to marginal-risk reasoning as such; it is to comparisons made “rather arbitrarily,” their specific charge against the widely-cited 10^26 FLOP compute threshold used in some governance proposals.

Their actual, broader critique — the throughline of that essay, restated by CASRAI’s own coverage — is that any single-model test, marginal or absolute, is a test of a model property, and safety “depends to a large extent on the context and the environment in which the AI model or AI system is deployed.” A well-evidenced marginal-risk comparison and a well-evidenced absolute-risk assessment share the same structural limit in their argument: both are still lab-run judgments about the model in isolation, not about the deployment context that actually determines whether harm occurs. In a later piece, “Do AI Risks Require Extraordinary Government Intervention?” (May 21, 2026), Kapoor and Narayanan argue for building societal resilience to AI misuse over precautionary restrictions on any single developer — a related but distinct argument from either risk-basis framing, and one that does not use “marginal” or “absolute” risk terminology at all. CASRAI found no essay by either author that argues for “absolute risk” as a framing developers should prefer; the accurate summary is narrower than that, and more interesting: they trust marginal-risk comparisons when the evidence behind them is rigorous, and doubt that any model-level test — whichever basis it uses — can substitute for evaluating the deployment context itself.

NIKOLAI’s shadow mapping of this exact framing

NIKOLAI is CASRAI’s own independent, unendorsed reference dictionary for frontier-AI-safety terminology — not a standard, and nothing here should be read as a mapping any named organization has reviewed or agreed to. Its N4 track, “Claims and Argument,” proposes an element for exactly the concept this guide describes: Marginal versus Absolute Risk Basis — verified live on this site on September 20, 2026, currently at nikolai-v0.1 and marked “proposed” status. NIKOLAI’s own operational definition names it as “a controlled value naming the baseline against which a stated risk is measured — e.g. a no-AI world, publicly available tools as of a stated date, other developers’ current models, or an industry-wide hypothetical where all developers matched the reporting developer’s practices,” added specifically because, in NIKOLAI’s own words, “at least four incompatible baselines are in active use across the corpus with no shared vocabulary connecting them” — the same four this guide documents above from Anthropic, OpenAI, Google DeepMind, and Meta’s own published text.

The element’s crosswalk carries five shadow-mapped rows — Anthropic, OpenAI, Google DeepMind, Meta, and the Frontier Model Forum — and every row is CASRAI’s own reading of each organization’s published document, not a mapping any of those organizations has declared, confirmed, or endorsed through NIKOLAI’s Mapping Declaration process. Two of the five (Anthropic and the Frontier Model Forum) are marked an exact match to NIKOLAI’s proposed definition; OpenAI, Google DeepMind, and Meta are marked close-but-not-equivalent, because each frames its own comparison slightly differently even where the underlying reasoning matches. NIKOLAI’s own divergence note is explicit that this is a live disagreement, not a settled vocabulary: “NIKOLAI’s crosswalk keeps these four baselines as distinct controlled values rather than treating ‘marginal risk’ as one interoperable term.” This guide treats that element as the shadow mapping it is — a proposal for future crosswalk material, not a declared equivalence between what these five organizations actually mean by “marginal risk.”

Frequently Asked Questions

Is “marginal risk” the same thing across Anthropic, OpenAI, Google DeepMind, and Meta?

No. All four reason about risk relative to some baseline rather than in isolation, but the baseline differs: Anthropic and OpenAI’s marginal-risk adjustment compare to other developers’ current models; OpenAI’s specific capability thresholds and Google DeepMind’s risk-acceptance test each separately compare to a fixed historical baseline (2021 tools, or a world without generative AI, respectively); and Meta’s “net new” test asks only whether general-purpose AI made an outcome newly possible at all. NIKOLAI’s own crosswalk of the concept documents this as four distinct, non-interoperable baselines rather than one shared term.

Did Narayanan and Kapoor argue that labs should use absolute risk instead of marginal risk?

No, and this guide corrects a common misreading. In “AI Safety Is Not a Model Property,” they cite a Stanford framework’s marginal-risk methodology approvingly, as the more rigorous approach to the specific question of open-model release. Their objection is to comparisons made “rather arbitrarily” without real evidence, and their broader argument — that a model-property test, whatever basis it uses, cannot by itself determine deployment safety — applies equally to marginal and absolute framings, not to one over the other.

Why does Anthropic’s RSP require Board approval specifically for marginal-risk decisions?

Because the policy treats marginal-risk reasoning as carrying its own risk: a model could be justified as “safer than the alternative” relative to competitors while still posing significant absolute danger. RSP v3.4 §3.4.5 requires that when marginal-risk analysis “plays a major role in a decision to move forward,” the Risk Report needs explicit approval from Anthropic’s Board and Long-Term Benefit Trust, not just its CEO and Responsible Scaling Officer — a higher bar than the policy’s default sign-off chain.

What is “risk creep,” and who coined it?

It is the Frontier Model Forum’s term, from its Risk Taxonomy and Thresholds report (§3.2), for the scenario where several developers each introduce a small marginal risk increase that individually clears a “net new” threshold, but which collectively raise industry-wide risk over time because no single release is ever compared against the cumulative total.

Is this the same question as CASRAI’s existing coverage of safety cases?

Related but distinct. CASRAI’s safety-case guide covers the argumentative structure — Claim, evidence, and the CAE (Claims, Arguments, Evidence) method GovAI has published on — that a developer uses to justify a single deployment decision. This guide covers a narrower, specific question inside that structure: which baseline a risk claim is measured against. A safety case built on a marginal-risk claim and one built on an absolute-risk claim can use the identical CAE structure while reaching very different conclusions about the same model.

Related CASRAI coverage

See also: Responsible Scaling Policy: what it is and how the major labs compare, Safety Cases: the missing methodology in frontier AI governance, the Narayanan-Kapoor critique of model-property safety testing, Bengio-Hinton vs. Narayanan-Kapoor on AI risk, capability thresholds across labs and regulators, and NIKOLAI’s Track System: A Map of the Frontier AI Safety Landscape (N1-N10).

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Marginal Risk vs. Absolute Risk: The Contested Framing Behind Frontier AI Safety Cases

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.