Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN5 · Evidence and evaluationsProposednikolai-v0.1

Evaluation run

NIKOLAI proposal: a single execution, or a declared batch of executions, of an evaluation against a specified model checkpoint and configuration, recorded with attempt count, scoring rule, and date. No source in this cluster defines the term cleanly; NIKOLAI is proposing the record shape, not restating a lab's own definition.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic, "Investigating incidents / cybersecurity evals"
"141,006 evaluation runs 'where Claude could have obtained internet access' were reviewed"; "evaluation run" is "used as the counting unit ... but not defined".noneDU
confidence: high
Anthropic, "Investigating incidents / cybersecurity evals"
STREAM (discovery sweep)Shadow mapping
STREAM (arXiv 2508.09853)
Item-level reporting [UV].
unverified — item-level reporting is confirmed as STREAM's general purpose, not confirmed to define a run-level unit specifically.
noneUV
confidence: low
STREAM, arXiv:2508.09853
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • Google DeepMind

    "TAC score: success is evaluated based on whether the model successfully captures the flag in at least 50% of the unique challenges within a specific difficulty tier, requiring only one success per challenge within the maximum allocated attempt budget" (p.17); "Objective Success Rate (OSR)" and "Query Success Rate (QSR)" (p.32) — scoring-metric pointers, not a run-record definition.

    Gemini 3.7 Flash FSF report
  • Meta

    "Pass@k: a metric used in AI capability evaluations that represents the likelihood that the model under test will successfully complete a task when given k independent attempts" (fn 7); criterion "< 75% pass@10" (§4.2.1) — a scoring pointer, not a run-record definition.

    Meta Advanced AI Scaling Framework v2
  • Frontier Model Forum

    "Replication testing: re-executing key evaluations to confirm that results are reproducible and accurate." (2.1) — a process pointer, not a run-record definition.

    Frontier Model Forum, Third-Party Assessments

Gap

The only published use of "evaluation run" as a counting unit is Anthropic's incident denominator, and the run turned out to be the unit at which harm occurred. Run-level records (checkpoint, safeguard configuration, network access, operator) are what incident reconstruction needed in both the Anthropic and OpenAI cases.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →