Source of record
Where this definition comes from
OpenAI Preparedness Framework v2, §3.1
“Scalable Evaluations: automated evaluations designed to measure proxies that approximate whether a capability threshold has been crossed.”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdfGoogle DeepMind Frontier Safety Framework v3.1, Glossary
“Early Warning Evaluations: are evaluations which measure the dangerous capabilities of a model. They specifically target the threats and risk scenarios identified through our threat modeling to determine a model's proximity to a CCL or TCL.”
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdfEU GPAI Code of Practice, Safety and Security Chapter, Measure 3.2
“Signatories will conduct at least state-of-the-art model evaluations in the modalities relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects.”
https://ec.europa.eu/newsroom/dae/redirection/document/118119
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match & verification | Source |
|---|---|---|---|
| AnthropicShadow mapping Anthropic Risk Report (August 2026) | “CoBench: "an internal evaluation measuring how well a model, placed at a historical point in Anthropic's infrastructure ..., can diagnose the root causes of issues that Anthropic engineers actually solved" (§3.4.3).” | exactEQ confidence: high | Anthropic Risk Report (August 2026) |
| AnthropicShadow mapping Anthropic Advanced AI Framework | “the safety framework must describe "capability evaluations performed" (p.5).” | exactEQ confidence: high | Anthropic Advanced AI Framework |
| OpenAIShadow mapping OpenAI Preparedness Framework v2 | “"Scalable Evaluations: automated evaluations designed to measure proxies that approximate whether a capability threshold has been crossed." "Deep Dives: designed to provide additional evidence validating the scalable evaluations' findings ..." (§3.1)” | exactEQ confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMindShadow mapping Frontier Safety Framework v3.1 | “"Early Warning Evaluations: are evaluations which measure the dangerous capabilities of a model. They specifically target the threats and risk scenarios identified through our threat modeling to determine a model's proximity to a CCL or TCL" (glossary).” | exactEQ confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAIShadow mapping xAI Frontier AI Framework (30 Jun 2026) | “"Model evaluations: state-of-the-art model evaluations relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects, which may include Q&A sets, task-based evaluations, benchmarks, red-teaming and other methods of adversarial testing, human uplift studies, model organisms, simulations, and/or proxy evaluations for classified materials" (s.2.2(2)).” CORRECTION: this row cites {FAIF26} but the original extraction omitted the required draft-document caveat. Adding it now: this document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final status was found (also unresolved per ALIGNMENT-MATRIX.md §6 item 4). Treat as provisional. | exactEQ confidence: medium | xAI Frontier AI Framework (30 Jun 2026) |
| xAIShadow mapping Grok 4.6 model card | “"safety-threshold evaluations"” | exactEQ confidence: high | Grok 4.6 model card |
| MetaShadow mapping Meta Advanced AI Scaling Framework v2 | “"Evaluation(s): refers to the assessments we do to understand capabilities and performance. We use this term to describe automated and human evaluations that assess capabilities, as well as evaluations to assess potential for misuse, such as red teaming and uplift studies." (Appendix I)” | exactEQ confidence: high | Meta Advanced AI Scaling Framework v2 |
| EUShadow mapping EU GPAI Code of Practice, Safety and Security Chapter | “Measure 3.2: "Signatories will conduct at least state-of-the-art model evaluations in the modalities relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects", including "open-ended testing ... with a view to identifying unexpected behaviours, capability boundaries, or emergent properties"; Appendix 3.1 requires "internal validity", "external validity" and "reproducibility".” | exactEQ confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
| EU / xAI (cross-reference)Shadow mapping EU CoP Appendix 3 example methods, echoed in xAI FAIF s.2.2(2) | “example methods (Appendix 3) "echoed by xAI s.2.2(2) almost verbatim": Q&A sets, task-based evaluations, benchmarks, red-teaming, human uplift studies, model organisms, simulations, proxy evaluations for classified materials.” This document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement was found disambiguating draft vs. final status as of this pass. Treat as provisional. | exactEQ confidence: medium | xAI Frontier AI Framework (30 Jun 2026) |
| California SB 53Shadow mapping California SB 53 | “Large developers summarise "(A) Assessments of catastrophic risks from the frontier model conducted pursuant to the large frontier developer's frontier AI framework" (22757.12(c)(2)).” | broadBR confidence: medium | California SB 53 |
| US Government (NIST CAISI)Shadow mapping NIST CAISI bulletin | “CAISI "will conduct pre-deployment evaluations and targeted research" (para. 1) — undefined.” | noneDU confidence: medium | NIST CAISI bulletin |
| Frontier Model ForumShadow mapping FMF Third-Party Assessments technical report | “"Capability Assessments: evaluate whether a model crosses any enabling capability thresholds or outcomes-based thresholds" (1.2).” | closeCL confidence: medium | Frontier Model Forum, Third-Party Assessments |
| Safety Framework Cards (discovery)Shadow mapping Safety Framework Cards (SSRN 7061798) | “"evaluation methodology" dimension [UV] — abstract-only, full text paywalled/unread.” unverified: full paper is SSRN account-gated; per the match-code table UV rows carry unverified:true. Open per ALIGNMENT-MATRIX.md §6 item 5. | noneUV confidence: low | Safety Framework Cards (SSRN 7061798, abstract only) |
| STREAM (discovery sweep)Shadow mapping STREAM (arXiv 2508.09853) | “"the only published item-level reporting standard for evaluations", with a template and gold-standard examples (discovery description) [UV].” unverified: only the abstract/scope has been confirmed from the primary source; the template's exact field names have not been read (ALIGNMENT-MATRIX.md §6 item 6). | noneUV confidence: low | STREAM, arXiv:2508.09853 |
What do these codes mean?
- exact
- The source term is equivalent to this element
- close
- The source term is close but not equivalent to this element
- broad
- The source term is broader than this element
- narrow
- The source term is narrower than this element
- none
- No mapping claim — used for false-friend and declared-but-undefined rows
- EQ
- Equivalent
- CL
- Close
- BR
- Source is broader than the element
- NR
- Source is narrower than the element
- FF
- False friend — same or similar label, different meaning
- DU
- Declared but undefined by the source
- UV
- Unverified
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- METR
"Full Capability Elicitation During Evaluations"; "Timing and Frequency of Evaluations" — named as process elements, not an evaluation definition (RL, not a mapping).
METR







