Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §2.16, §2.16.1
“Explicit limitations section (§2.16); "Evaluation awareness is conceded to 'partially undermine' confidence in the alignment assessment"; "we have not provided clear evidence that this elicitation is sufficiently strong" (§2.16.1).”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfOpenAI Preparedness Framework v2, §4.2
“Safeguards Report contents include "Any notable limitations with the information provided" (§4.2).”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match & verification | Source |
|---|---|---|---|
| AnthropicShadow mapping Anthropic Risk Report, August 2026 | “Explicit limitations section (§2.16); "Evaluation awareness is conceded to 'partially undermine' confidence in the alignment assessment"; "we have not provided clear evidence that this elicitation is sufficiently strong" (§2.16.1).” | exactEQ confidence: high | Anthropic Risk Report, August 2026 |
| OpenAIShadow mapping OpenAI Preparedness Framework v2 | “Safeguards Report contents include "Any notable limitations with the information provided" (§4.2).” | exactEQ confidence: high | OpenAI Preparedness Framework v2 |
| MetaShadow mapping Meta Advanced AI Scaling Framework v2 | “"Preparedness reports will also disclose any known issues that could hinder generalizing our safety testing to realworld risks, including changes to the training process that reduce interpretability" (§2.2.1).” | exactEQ confidence: high | Meta Advanced AI Scaling Framework v2 |
| California SB 53Shadow mapping California SB 53 | “Transparency report "(G) Any generally applicable restrictions or conditions on uses of the frontier model" (22757.12(c)(1)). This is a use restriction on the deployed model, not a stated limitation of the assessment's conclusions.” FALSE FRIEND, explicitly coded FF in the source table: SB 53's transparency-report item labelled 'restrictions or conditions on uses' is a product/usage-policy disclosure, not an assessment-limitation disclosure. See the element-level divergence_note. | noneFF confidence: high | California SB 53 |
| STREAM (discovery sweep)Shadow mapping Discovery sweep (evaluator ecosystem) | “Evaluation reporting template (content not read at time of the working table) [UV].” STREAM's evaluation reporting template (arXiv 2508.09853) was subsequently confirmed open-access and fetched directly (ALIGNMENT-MATRIX.md §6 item 6, resolved fourth pass 16 Sep 2026): title, scope (a pilot limited to ChemBio benchmarks, 23 contributing experts), format (three-page template with worked 'gold standard' examples), and purpose (comparable, item-level evaluation disclosure) are confirmed from the primary source. Still open: the template's exact field names -- needed to know whether it has a distinct 'limitations' field matching this element -- require reading the PDF body, not yet done. Match type left as 'none' rather than guessed; citation URL corrected to the resolved arXiv page (the sources table's {DEV} tag names 'STREAM=arXiv 2508.09853' but has no single URL of its own, so the arXiv abstract-page URL is used here rather than the compound {DEV} tag). | noneUV confidence: low | STREAM evaluation reporting template (arXiv 2508.09853; per {DEV} sources-table entry) |
| AnthropicShadow mapping Anthropic Advanced AI Framework | “AAF system card contents: "Model capabilities and limitations, as well as intended and observed model behaviors" (p.6).” | closeCL confidence: high | Anthropic Advanced AI Framework |
What do these codes mean?
- exact
- The source term is equivalent to this element
- close
- The source term is close but not equivalent to this element
- broad
- The source term is broader than this element
- narrow
- The source term is narrower than this element
- none
- No mapping claim — used for false-friend and declared-but-undefined rows
- EQ
- Equivalent
- CL
- Close
- BR
- Source is broader than the element
- NR
- Source is narrower than the element
- FF
- False friend — same or similar label, different meaning
- DU
- Declared but undefined by the source
- UV
- Unverified
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Google DeepMind
RL -- GDM's report discusses a limitation-adjacent concern but the source explicitly records that GDM has no distinct limitation element/field; recorded as a pointer, not a mapping.
Gemini 3.7 Flash FSF Report - METR
RL -- a specific investigation's own limitations, not a general Limitation record type or field definition.
METR OpenAI/Hugging Face Incident Investigation blog - Frontier Model Forum
RL -- a reporting convention within FMF's third-party assessment sub-types, not a standalone limitation definition.
Frontier Model Forum: Third-Party Assessments
Divergence
Where sources materially disagree
REQUIRED false-friend note (an SB 53 crosswalk row above is coded FF): California SB 53's transparency-report item '(G) Any generally applicable restrictions or conditions on uses of the frontier model' (22757.12(c)(1)) uses limitation-adjacent language but names a product/usage-policy restriction imposed BY the developer ON users of the deployed model -- e.g. acceptable-use terms -- not a stated reason why the developer's own safety assessment might not hold or generalise, which is what every other source in this cluster means by 'limitation'. A naive string or keyword match on 'restrictions'/'limitations' between SB 53 and the lab frameworks would wrongly equate a use-policy disclosure with an assessment-caveat disclosure; NIKOLAI's crosswalk keeps them separate. Distinguish from evaluation-validity-threat (N5): evaluation-validity-threat is the CONTROLLED-VALUE tag attached to a specific evaluation-run record (sandbagging, evaluation awareness, alignment faking, metagaming/grader-gaming, reward hacking); limitation is the free-text caveat attached to the resulting CLAIM or safety-case. A limitation arising from an evaluation-validity threat should reference the corresponding evaluation-validity-threat value rather than restating it in prose -- e.g. "this claim's confidence is limited by [evaluation awareness]", not a freestanding description of evaluation awareness with no link back.







