Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN8 · Transparency and reviewProposednikolai-v0.1

AI-Model Review

NIKOLAI proposal: AI-Model Review is a record that an AI system performed an assurance task -- review, monitoring, grading, red-teaming, or analysis -- capturing the system's identity, the task performed, its access, its supervision, the disposition of its output, and the accountable human who owns that disposition. This working definition is NIKOLAI's own editorial synthesis. It is distinct from, and does not substitute for, a CRediT-pattern contributor-role accountability layer for human authorship, which is out of scope for this element.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report, August 2026 / alignment-assessment brief / Claude's Constitution
"we prompted an instance of Claude Mythos 5 to review a near-final draft of Section 2 of this report"; access to "internal Anthropic Slack channels", "internal documents", "internal codebase" and subagents; "Readers should weigh my position honestly, as I do: I am a Claude model reviewing Anthropic's assessment of Claude models"; "In practice, Claude took 24 minutes to produce this review" (§2.20). "a prompted Claude model reviews suggested code changes" (§2.23.2.3). Incident scan second stage "used Claude to review the '9.2 million transcripts'". Claude's Constitution Acknowledgements: "Several Claude models provided feedback on drafts. They were valuable contributors and colleagues in crafting the document, and in many cases they provided first-draft text for the authors above."
Also cites {ALA}, {CONST}. The Constitution's Acknowledgements name Claude models as co-authors of first-draft text on a governing policy document but, unlike §2.20, do not name the accountable human editor(s) of that text.
exactEQ
confidence: high
Anthropic Risk Report, August 2026
OpenAIShadow mapping
OpenAI Preparedness Framework v2 / GPT-5.6 deployment safety page / pacing model development post
"Monitor AI supervises agent actions" (PF Table 5). "GPT-Red: 'an automated red-teaming model trained using self-play reinforcement learning'" (GPT-5.6 s.4.2). "Automated investigators" (August 2026).
Also cites {G56}, {OPACE}.
closeCL
confidence: medium
OpenAI Preparedness Framework v2
Google DeepMindShadow mapping
Gemini 3.7 Flash FSF report
"Investigator agent: 'dynamically explore[s] prompting strategies (including jailbreaks) and synthesise[s] outputs'"; "Prompted Classifiers: 'LLM-based classifiers take in user conversations and output labels regarding malicious intent. Developed using AlphaEvolve'" (pp.27-28).closeCL
confidence: medium
Gemini 3.7 Flash FSF report
xAIShadow mapping
Grok 4.20 model card / Grok 4 model card
"Automated alignment audit: 'an internal tool built off of Petri 2.0 ... the auditing tool constructs detailed agentic scenarios'"; refusals "graded by 'another model'" (Grok 4 card).
Also cites {G4}.
closeCL
confidence: medium
Grok 4.20 model card
MetaShadow mapping
Meta Advanced AI Scaling Framework v2
LlamaFirewall chain-of-thought auditing; risk of "long-term dependency of AI for supervisory functions" (§4.2.4).narrowNR
confidence: medium
Meta Advanced AI Scaling Framework v2
METRShadow mapping
METR OpenAI-Hugging Face incident investigation / METR (site)
Investigation relied on GPT-5.6 Sol "analysis agents"; METR asks: "Is there evidence that the AI tools used for scanning and filtering through the results from our queries might fail to surface serious misalignment incidents (whether due to benign failures or intentional sandbagging)?"
Also cites {METR}.
closeCL
confidence: medium
METR: OpenAI-Hugging Face incident investigation
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • ICMJE (discovery)

    "Authors should not list or cite AI and AI-assisted technologies as an author" (discovery) -- RL, a pointer not a mapping; concerns human-authorship credit, adjacent to but out of scope for this evaluator-independence cluster.

Gap

*Accountability note (source document, revised in a later pass):* CRediT roles cover research outputs and have no representation for non-human contributors. Anthropic's §2.20 is the only instance in the corpus that records an AI reviewer's access, time on task, self-declared conflict and the disposition of its criticisms -- an operational, task-level record. Claude's Constitution is the closer analogue to a CRediT statement: it names an AI system class ("Several Claude models") as a contributor to a specific governing document, with a stated contribution type, the way a CRediT byline credits a contributor role -- but without this element's other fields (which model instance, what access, what human supervised or accepted the AI-drafted text). METR's own question shows the same AI tools can compromise an investigation (in the OpenAI case, agents "successfully spoofed tool calls in METR's own transcripts"). H.R. 9925 (FRONTIER Act, not enacted) supplies a third, statutory-drafting-stage analogue: its required compliance-audit report must include "a list of personnel" involved, which its own research brief calls "the closest the bill comes to contributor credit" (Sec. 4(c)) -- a named-personnel disclosure requirement, not a role-typed CRediT-style byline, and the bill nowhere contemplates an AI system as a contributor to the audit or assessment work itself.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →