Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Monitor

NIKOLAI proposes to define a Monitor record as: an automated or human process that observes model inputs, outputs, reasoning, actions, or internal state to detect a specified behavior, carrying a stated coverage scope, sampling rate, and escalation path when a detection fires. This is an unsourced NIKOLAI editorial synthesis of how Anthropic, OpenAI, and Google DeepMind each independently operate chain-of-thought and output monitors as a distinct mitigation category from access-control safeguards.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report, August 2026 / Training a Misaligned Reward Seeker
an Opus 4.6 completion-only monitor (no subsampling for the most common agent scaffolds; 5% for tool-marked completions), an Opus 4.8 prompt+completion monitor on a 25% subsample of impactful-tool-call traffic ... surface ~50 trajectories per week for manual review by AI control researchers; "we monitor the Chain-of-Thought to detect and measure reward hacking, but importantly do not reward or penalize the model based on it, in order to preserve monitorability"
Term also draws on {RSEEK} ("Training a Misaligned Reward Seeker", 16 Sep 2026 correction pass) for the CoT-monitoring/severity-flagging quote.
exactEQ
confidence: high
Anthropic Risk Report, August 2026
OpenAIShadow mapping
OpenAI Preparedness Framework v2 / Pacing Model Development (Cyber), Aug 2026
"Monitor AI supervises agent actions"; "Monitoring (safeguard): 'which detects and allows us to respond to concerning behavior'"; "Activation classifiers: 'detectors that inspect a model's internal activity'"; "Automated investigators"
Also cites {OPACE} (Aug 2026 post) for activation classifiers / automated investigators language.
exactEQ
confidence: high
OpenAI Preparedness Framework v2
Google DeepMindShadow mapping
Frontier Safety Framework v3.1
Deployment mitigations include "input/output/chain-of-thought monitoring and analysis" (glossary); safeguards "such as chain-of-thought monitoring on high-risk deployments"closeCL
confidence: medium
Google DeepMind Frontier Safety Framework v3.1
xAIShadow mapping
Frontier AI Framework, 30 Jun 2026
Incident detection "Monitoring and alerting of public comments from the X platform"; security "real-time monitoring"
xAI's Frontier AI Framework (30 Jun 2026) carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat this citation as provisional.
narrowNR
confidence: medium
xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata)
MetaShadow mapping
Meta Advanced AI Scaling Framework v2
LlamaFirewall "chain-of-thought auditing" and input classifiersnarrowNR
confidence: medium
Meta Advanced AI Scaling Framework v2
EUShadow mapping
EU GPAI Code of Practice, Safety and Security Chapter
Measure 3.5 "Post-market monitoring": Signatories "will conduct appropriate post-market monitoring to gather information relevant to assessing whether the systemic risk could be determined to not be acceptable ... and to inform whether a Model Report update is necessary", via methods including end-user feedback, incident reporting forms, bug bounties, "monitoring software repositories, known malware, public forums, and/or social media", and "monitoring aspects of models that are relevant for assessing and mitigating systemic risk and are not transparent to third parties, e.g. hidden chains-of-thought"
Source notes this confirms xAI's chapter-derived s.2.2(5) wording; the linked {FAIF26} carries the draft/final metadata caveat above.
exactEQ
confidence: high
EU GPAI Code of Practice, Safety and Security Chapter
Frontier Model ForumShadow mapping
FMF Information Sharing / Incident Reporting Issue Brief
"Monitoring and Detection Systems: Enhancing systems that detect anomalous behavior, unauthorized access, or potential misuse, for safety and security purposes only" (Table 3)closeCL
confidence: medium
Frontier Model Forum, Information Sharing / Incident Reporting Issue Brief
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • SB 53 (California)

    Framework topic "(10) ... including risks resulting from a frontier model circumventing oversight mechanisms" (22757.12(a)) — related to monitoring but a pointer, not a mapping (RL).

    California SB 53
  • UK AISI / Google DeepMind

    "CoT monitoring helps us understand how an AI system produces its answers, complementing interpretability research." — related, not a direct monitor-record mapping (RL).

    DeepMind, Deepening Our Partnership with UK AI Security Institute

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →