Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §1.3.1, §1.3.2, §3.2
“"AI R&D threshold (current)": met if "(1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is 'dramatic acceleration' of the pace of AI progress for reasons that likely relate to the automation of AI R&D" (§1.3.1, §3.2).”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfGoogle DeepMind Frontier Safety Framework v3.1, glossary, p.18
“"Critical Capability Levels (CCLs): are the main capability thresholds around which we have built the Framework process." "Tracked Capability Levels (TCLs): are capability thresholds which capture a lower level of risks than our CCLs." (glossary, p.18)”
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdfMagic AGI Readiness Policy, Threshold Definition
“The sole disclosed numeric public capability threshold in the corpus — "when, at the end of a training run, our models exceed a threshold of 50% accuracy on LiveCodeBench" (Pass@1), with the contemporaneous public-model baseline "Claude-3.5-Sonnet 48.8%" given for comparison.”
https://magic.dev/agi-readiness-policy
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match & verification | Source |
|---|---|---|---|
| AnthropicShadow mapping Anthropic Risk Report, August 2026 / RSP v3.4 | “"AI R&D threshold (current)": met if "(1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is 'dramatic acceleration' of the pace of AI progress for reasons that likely relate to the automation of AI R&D" (§1.3.1, §3.2). CB-2 threshold: "AI systems that can functionally substitute for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons with potential for catastrophic harm ..." (§1.3.2). RSP §1 table header: "Capability or usage threshold."” Source document's own citation for the RSP §1 table-header note also carried a second bracket tag ({DAN}) that does not resolve against NIKOLAI's provided sources table; omitted here rather than guessed at. | exactEQ confidence: high | Anthropic Risk Report, August 2026 |
| OpenAIShadow mapping Preparedness Framework v2 / Frontier Governance Framework | “"Capability thresholds: concretely describe things an AI system might be able to help someone do or might be able to do on its own that could meaningfully increase risk of severe harm." (§2.2). FGF: capability threshold is "used but never formally defined"; "Risk tiers" are the published scale (§2.2, §2.3).” | exactEQ confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMindShadow mapping Frontier Safety Framework v3.1 | “"Critical Capability Levels (CCLs): are the main capability thresholds around which we have built the Framework process." "Tracked Capability Levels (TCLs): are capability thresholds which capture a lower level of risks than our CCLs." (glossary, p.18)” | exactEQ confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAIShadow mapping xAI Frontier AI Framework (30 Jun 2026) / Grok 4.6 card / earlier xAI frameworks | “June 2026 framework: "When analyzing whether a threshold has been reached ..." (s.2.3), with no thresholds stated. Grok 4.6 card: "FAIF safety thresholds" cited with no values given. The earlier Dec 2025 FAIF and Aug 2025 RMF used deployment acceptance criteria instead: "maintaining an answer rate of less than 1 out of 20 on restricted queries" and "a dishonesty rate of less than 1 out of 2 on MASK."” PDF metadata /Title of the 30 Jun 2026 framework reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final was found. | noneDU confidence: medium | xAI Frontier AI Framework (30 Jun 2026, draft-marked) |
| EUShadow mapping EU GPAI Code of Practice, Safety and Security chapter | “Measure 4.1: for each identified systemic risk, Signatories "define appropriate systemic risk tiers that: (i) are defined in terms of model capabilities, and may additionally incorporate model propensities, risk estimates, and/or other suitable metrics; (ii) are measurable; and (iii) comprise at least one systemic risk tier that has not been reached by the model." Content of tiers left to each Signatory; no shared tier scale published in the chapter.” | exactEQ confidence: high | EU GPAI Code of Practice, Safety and Security chapter |
| California SB 53Shadow mapping California SB 53 | “Frameworks must describe "(2) Defining and assessing thresholds used by the large frontier developer to identify and assess whether a frontier model has capabilities that could pose a catastrophic risk, which may include multiple-tiered thresholds." (22757.12(a)). "Threshold" itself is not defined by the statute.” | noneDU confidence: medium | California SB 53 |
| US Government (Executive Order 14409)Shadow mapping EO 14409 / Seoul Frontier AI Safety Commitments | “EO 14409 designation threshold is classified (cyber domain). Seoul Commitment II: thresholds are "at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable" — a narrower, outcome-gated framing than a capability threshold proper.” | narrowNR confidence: medium | Executive Order 14409 |
| METRShadow mapping METR Common Elements | “"Capability Thresholds": "Thresholds at which specific AI capabilities would pose severe risk and require new mitigations."” | exactEQ confidence: high | METR (common-elements) |
| Frontier Model ForumShadow mapping FMF Risk Taxonomy and Thresholds | “"Enabling Capability Thresholds (also called 'critical capability levels' or 'capability thresholds'): Abilities that could potentially enable extreme harms if the model is deployed without additional safeguards." (s3.1, p.10)” | exactEQ confidence: high | FMF Risk Taxonomy and Thresholds |
| Safety Framework Cards (discovery)Shadow mapping Safety Framework Cards (SSRN 7061798) | “"capability thresholds" dimension named in discovery sweep; full text paywalled/unread.” Unverified: full text inaccessible to the source document's own research pass. | noneUV confidence: low | Discovery sweep — Safety Framework Cards (SSRN 7061798, paywalled/unread) |
| AmazonShadow mapping Amazon Frontier Model Safety Framework | “"Critical Capability Thresholds": "a set of model capabilities that have the potential to cause significant harm to the public if misused" (Overview, p.1); one qualitative threshold per domain (CBRN, Offensive Cyber Operations, Automated AI R&D), stated as an uplift-based description rather than a benchmark score.” | exactEQ confidence: high | Amazon Frontier Model Safety Framework |
| MagicShadow mapping Magic AGI Readiness Policy | “The sole disclosed numeric public capability threshold in the corpus — "when, at the end of a training run, our models exceed a threshold of 50% accuracy on LiveCodeBench" (Pass@1), with the contemporaneous public-model baseline "Claude-3.5-Sonnet 48.8%" given for comparison (Threshold Definition).” | exactEQ confidence: high | Magic AGI Readiness Policy |
What do these codes mean?
- exact
- The source term is equivalent to this element
- close
- The source term is close but not equivalent to this element
- broad
- The source term is broader than this element
- narrow
- The source term is narrower than this element
- none
- No mapping claim — used for false-friend and declared-but-undefined rows
- EQ
- Equivalent
- CL
- Close
- BR
- Source is broader than the element
- NR
- Source is narrower than the element
- FF
- False friend — same or similar label, different meaning
- DU
- Declared but undefined by the source
- UV
- Unverified
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Meta
"Risk thresholds: are the incremental levels of risk that a Frontier AI model might pose towards realization of a catastrophic outcome" (Appendix I) — these are outcome-based risk levels, not capability thresholds; capability enters only through separately-named "Enabling capabilities" and "Operational threshold" (cyber). Scored RL (related, not a mapping) in the source document.
Meta Advanced AI Scaling Framework v2
Gap
SB 53 requires thresholds but defines none; xAI's current framework and the Grok 4.6 card cite thresholds that are not published; Amazon's thresholds are qualitative only, with undefined terms such as "material uplift" and "reliably." Magic is the one exception with a disclosed numeric public threshold. NIKOLAI needs a `threshold.disclosureStatus` value set (quantified; qualitative; referenced-but-undefined; classified).







