Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN3 · Thresholds and checkpointsProposednikolai-v0.1

Threshold Status

NIKOLAI editorial proposal (unsourced): threshold status is the recorded determination of a model's position relative to a capability threshold (not reached; cannot rule out; provisionally met; met), together with the date, evidentiary basis and determining party. This is element B5 of the source crosswalk. A cross-lab controlled vocabulary appears to be converging: GDM's "cannot rule out" ≈ OpenAI's "precautionarily treated as reached" ≈ Anthropic's "provisionally met" ≈ Meta's "provisionally rated" — the same precautionary intermediate state under four different names, observed across four labs.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report, August 2026
Every frontier model since Opus 4 is treated as "provisionally meeting the CB-1 threshold" "to err on the side of caution" (§4.4.2). On AI R&D: "Neither RSP criterion is met" (§3.4).closeCL
confidence: high
Anthropic Risk Report, August 2026
OpenAIShadow mapping
Preparedness Framework v2 / GPT-5.6 deployment safety report
SAG options: threshold crossed; threshold not crossed; recommend a deep dive (§3.3). GPT-5.6 designations: "High in Biological and Chemical", "High in Cybersecurity", "below High in AI Self-Improvement"; "these models should thus be precautionarily treated as High" (s.9, s.9.1.1).exactEQ
confidence: high
OpenAI Preparedness Framework v2
Google DeepMindShadow mapping
Gemini 3.7 Flash FSF report / model card
"if we cannot rule out, based on the evidence and threat models we have, that a T/CCL has been reached, we designate the model as 'cannot rule out being at the T/CCL', and mitigate accordingly" (report p.4). Outcomes recorded as "No T/CCL reached" and "CBRN Uplift 1 CCL alert threshold reached" (pp.3, 12). Model card column "CCL reached?" with value "CCL not reached."exactEQ
confidence: high
Gemini 3.7 Flash FSF report
xAIShadow mapping
Grok 4.6 card / Grok 4.20 card
"Grok 4.6 scores below the FAIF safety thresholds on dual-use knowledge, indicating limited actionable uplift for an already-trained actor." (§8, p.31). Grok 4.20: released "with safeguards appropriate for its capability threshold" (s.1.3), with the threshold left unidentified.closeCL
confidence: medium
Grok 4.6 model card
MetaShadow mapping
Meta Advanced AI Scaling Framework v2
"Until evaluation on the complex suite of challenges is completed, any model meeting the simple-suite threshold is provisionally rated 'high' risk for the given deployment scenario" (§4.2.1, p.29). Assignment: "the Chief AI Officer and Director of Alignment and Risk will assign a risk threshold" (§2.1.2).closeCL
confidence: high
Meta Advanced AI Scaling Framework v2
EUShadow mapping
EU GPAI Code of Practice, Safety and Security chapter
Framework must require "at least one systemic risk tier that has not been reached by the model" (Measure 4.1); the acceptance determination itself is a binary "acceptable" / "not determined to be acceptable" outcome per identified risk and overall (Measure 4.1(2)-(3), Measure 4.2), rather than a named "cannot rule out" intermediate status.closeCL
confidence: high
EU GPAI Code of Practice, Safety and Security chapter
US Government (Executive Order 14409)Shadow mapping
EO 14409
Developers may "engage the Federal Government to determine whether model(s) under development meet the designation of 'covered frontier model'" (Sec. 3(b)(i)) — a narrower, government-facing determination step rather than a full status vocabulary.narrowNR
confidence: medium
Executive Order 14409
METRShadow mapping
METR Common Elements
Thresholds "are compared to the results of model evaluations to determine whether they have been crossed."closeCL
confidence: high
METR (common-elements)
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • Frontier Model Forum

    Crossing a threshold "signals entry into a new phase of heightened risk, where more rigorous risk assessments for this domain and stronger baseline safety and security measures are potentially warranted" (s3.1) — describes a consequence of status change, not a status vocabulary itself. Scored RL in the source document.

    FMF Risk Taxonomy and Thresholds

Gap

Controlled vocabulary emerging from the sources: not reached / alert threshold reached / cannot rule out (GDM) = precautionarily treated as reached (OpenAI) = provisionally met (Anthropic) = provisionally rated (Meta) / reached. The precautionary intermediate state is common to four labs under four different names — a strong candidate for a single NIKOLAI controlled-value ladder.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →