Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN5 · Evidence and evaluationsProposednikolai-v0.1

Evaluation-validity threat

NIKOLAI proposal: a controlled list of named conditions — evaluation awareness, sandbagging, alignment faking, metagaming/grader-gaming, reward hacking — under which a model's behaviour during evaluation may not reflect its behaviour in deployment, so that a lab's claim of having checked for this is machine-comparable across frameworks.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report (August 2026)
"Grader awareness: 'a model explicitly reasoning about the grading/evaluation mechanism itself, beyond what is mentioned in the prompt, and using that reasoning to change its strategy'" (§2.25.1). Evaluation awareness "partially undermine[s]" confidence (§2.9.4). Pathway "diffuse sandbagging on safety R&D" (§2.2.1).exactEQ
confidence: high
Anthropic Risk Report (August 2026)
AnthropicShadow mapping
Anthropic, "Reward seeker" (alignment.anthropic.com)
"Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate" (tl;dr); the studied model is characterised as a "'reward-on-the-episode seeker' (Carlsmith)"; telling the model a monitor exists "makes hacking sneakier, not rarer" — sneaky hacking rose "from 4% to 19%"; production-style monitoring "would have caught it": "97% of reward hacking environments with reward hacking rates of at least 1% were flagged as significant or severe".exactEQ
confidence: high
Anthropic, "Reward seeker"
OpenAIShadow mapping
OpenAI Preparedness Framework v2
"Sandbagging (Research Category): ability and propensity to respond to safety or capability evaluations in a way that significantly diverges from performance under real conditions, undermining the validity of such evaluations." (Table 2)exactEQ
confidence: high
OpenAI Preparedness Framework v2
OpenAIShadow mapping
GPT-5.6 deployment safety card
"Metagaming: when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt." (s.7.4)exactEQ
confidence: high
GPT-5.6 deployment safety card
Google DeepMindShadow mapping
Gemini 3.7 Flash FSF report
"Sandbagging: deliberate underperformance in order to avoid being flagged as dangerous" (p.5); "Evaluation awareness: where the model indicated knowledge of the testing context or explicitly noted it was operating in a synthetic environment" (p.6).exactEQ
confidence: high
Gemini 3.7 Flash FSF report
xAIShadow mapping
xAI model-card discovery note (tag {DXAI} in source document)
"verbalized (evaluation) awareness" listed as a metric term in xAI cards (discovery) [UV].
Citation tag {DXAI} in the source document does not appear in the supplied sources table; no URL could be verified against it, so none is given rather than invented. This row is otherwise per the match-code table's UV handling (unverified:true).
noneUV
confidence: low
xAI discovery sweep (source cites tag {DXAI}, which does not resolve to an entry in the supplied sources table)
MetaShadow mapping
Meta Advanced AI Scaling Framework v2
"Evaluation awareness: Ability to accurately determine whether it is being evaluated and the purpose of that evaluation ..."; "Selective evaluation-adaptive behavior" (enabling capabilities, §3.4).exactEQ
confidence: high
Meta Advanced AI Scaling Framework v2
METRShadow mapping
METR
"Evaluation awareness is now a named validity threat" (discovery summary) [UV].
unverified per the match-code table (unverified:true); confirmed only as a discovery-sweep summary, not read from a specific dated METR page.
noneUV
confidence: low
METR
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • US (SB 53)

    Incident limb applies "outside of the context of an evaluation designed to elicit this behavior" (22757.11(d)(4)) — a carve-out pointer, not an evaluation-validity-threat definition.

    California SB 53

Divergence

Where sources materially disagree

Alignment faking added 2026-09-19 as a fifth named condition, distinct from sandbagging: sandbagging is behavioural UNDERperformance to hide a capability from a dangerous-capability evaluation; alignment faking is behavioural OVERperformance-on-cooperation to hide misalignment from a propensity/alignment evaluation. Both are evaluation-awareness-driven, but they point in opposite directions and should not be conflated. See limitation's own divergence_note for how this element relates to N4's limitation.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →