Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Robustness level

NIKOLAI proposes Robustness level as a graded controlled-value axis stating how hard a safeguard is to circumvent, paired with the attack model assumed and the evidence supporting the grade (e.g. red-team success rate, independently verified multiplier). This is an unsourced NIKOLAI editorial synthesis; Anthropic and Google DeepMind both maintain a named, graded robustness/adversarial-robustness axis, which supports treating it as a distinct controlled-value ladder from coverage level.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. A shadow row is CASRAI's own reading of a published document. No lab, evaluator or regulator named on a shadow row has declared, endorsed, or been consulted on it. That changes only when an organisation files its own Mapping Declaration.
OrganisationTheir term, as publishedMatch & verificationSource
AnthropicShadow mapping
Anthropic Risk Report, August 2026
Level 1 = the design from the previous Risk Report, patched for known jailbreaks; Level 2 = updated Constitutional Classifiers (Jan 2026 paper), streaming linear probe + fine-tuned second stage, weighted combination; Level 3 = Level 2 "with a lower threshold for blocking queries in order to be more conservative for our most capable models"exactEQ
confidence: high
Anthropic Risk Report, August 2026
OpenAIShadow mapping
OpenAI Preparedness Framework v2 / GPT-5.6
"Robustness (claim): Users cannot use the model to cause the harm because they cannot elicit the capability, such as because the model is modified to refuse to provide assistance to harmful tasks and is robust to jailbreaks that would circumvent those refusals." Efficacy metrics include "Time to patching a new known jailbreak"; GPT-5.6 "Worst-case defender success rate"closeCL
confidence: medium
OpenAI Preparedness Framework v2
Google DeepMindShadow mapping
Gemini 3.7 Flash FSF Report
"Adversarial Robustness: 'if threat actors do try to circumvent safeguards on "covered" topics, how well would they be able to do so?'"; "Violation Rate"closeCL
confidence: medium
Gemini 3.7 Flash FSF Report
MetaShadow mapping
Meta Advanced AI Scaling Framework v2
CB acceptance: "40% refusal or safe responses against all adversarial attacks within a typical adversarial attack portfolio"narrowNR
confidence: medium
Meta Advanced AI Scaling Framework v2
G42Shadow mapping
G42 Frontier Safety Framework
"DML 2" objective: "Even a determined actor should not be able to reliably elicit CBRN weapons advice ..."
Source document flags this row assignment as an open [VERIFY] item (ALIGNMENT-MATRIX.md §6 item 16, "G42 DML 2 row assignment" — not resolved as of the 16 Sep 2026 pass); treat the row placement itself, not just the content, as provisional.
closeCL
confidence: low
G42 Frontier Safety Framework
What do these codes mean?
exact
The source term is equivalent to this element
close
The source term is close but not equivalent to this element
broad
The source term is broader than this element
narrow
The source term is narrower than this element
none
No mapping claim — used for false-friend and declared-but-undefined rows
EQ
Equivalent
CL
Close
BR
Source is broader than the element
NR
Source is narrower than the element
FF
False friend — same or similar label, different meaning
DU
Declared but undefined by the source
UV
Unverified

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • xAI

    HackerBench "under 'the standard release-tracked safeguards': 6.9% harmful/dual-use compliance and 0.0% benign refusal" — related benchmark result, not a mapping (RL).

    Grok 4.6 Model Card
  • UK AISI (via Anthropic)

    "Boundary Point Jailbreaking (BPJ)": "an automated methodology developed by UK AISI that optimizes against black-box classifiers"; AISI "independently verified the Level 3 robustness multiple (AISI estimated 2x, Anthropic 3x)" — related methodology/independent-verification note, not a mapping (RL).

    Anthropic Risk Report, August 2026
  • Frontier Model Forum

    "Adversarial testing: structured attempts to circumvent model safeguards and elicit harmful behaviors through techniques that malicious actors might employ." — related, not a mapping (RL).

    Frontier Model Forum, Third-Party Assessments

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →