Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN5 · Evidence and evaluationsProposednikolai-v0.1

Saturation status

NIKOLAI proposal: a controlled value recording whether an evaluation can still discriminate meaningfully at the relevant capability level, the criterion used to decide that, and the consequence (retire, replace, or treat as met). No shared numeric saturation criterion exists across sources — proposing a controlled ladder, rather than adopting one lab's number, is the standards contribution here.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report (August 2026)
"our most concrete task-based evaluations have 'saturated'" (§3.1); CB-1 evaluation deprioritised "because many have saturated" (§4.4.2).exactEQ
confidence: high
Anthropic Risk Report (August 2026)
AnthropicShadow mapping
Anthropic Institute, "Recursive self-improvement"
"Benchmarks measure the performance of models in a given domain, and they're 'saturated' when models achieve close to 100% performance."exactEQ
confidence: high
Anthropic Institute, "Recursive self-improvement"
OpenAIShadow mapping
GPT-5.6 deployment safety card
AI self-improvement evaluations "were replaced with a new suite ... because older ones were saturated or flawed" (s.9.1.3).closeCL
confidence: high
GPT-5.6 deployment safety card
Demis Hassabis (personal essay)Shadow mapping
Hassabis substack post
"These evaluations would be regularly updated, perhaps quarterly to start, with outdated or saturated benchmarks being deprecated and replaced."closeCL
confidence: medium
Demis Hassabis, substack post
MicrosoftShadow mapping
Microsoft Frontier Governance Framework
benchmark inclusion criterion: "have low saturation (i.e., the best performing models typically score lower than 70%)" (fn 1, p.5).closeCL
confidence: high
Microsoft Frontier Governance Framework
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • Meta

    Saturation used as an acceleration indicator: "the model saturates benchmarks in each domain six months or less after public release of a benchmark" (fn 8) — a pointer use of the term, not a saturation-status definition.

    Meta Advanced AI Scaling Framework v2

Gap

Anthropic defines saturation near 100% model performance; Microsoft sets a 70% ceiling for benchmark inclusion; Hassabis proposes quarterly deprecation of outdated/saturated benchmarks. A standards body refreshing benchmarks would need a shared saturation criterion; none currently exists.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →