Skip to main content
v2026.11,858 entries · CC-BY 4.0
Frontier AI Safety & Governance

Third-Party Evaluation & Assurance

The independent evaluation ecosystem around frontier AI: METR, the UK AI Security Institute, the US CAISI, embedded and arms-length evaluator models, evaluator independence standards, and third-party AI auditing.

Guides

ISO/IEC 42006 Explained: Who Is Allowed to Certify ISO 42001

ISO/IEC 42006:2025 is not a standard you get certified against — it is the criteria document accreditation bodies use to assess the certification bodies that issue ISO/IEC 42001 certificates. Published 7 July 2025, nineteen months after ISO 42001 itself. What it is, the three-layer stack behind every certificate, and a five-minute register check that tells you whether a vendor certificate means anything.

Irregular and the “Frontier Security Lab” Category: What Cyber-Capability Evaluation Actually Involves

Irregular (formerly Pattern Labs) coined the term “frontier security lab” for itself in September 2025. Its SOLVE, CyScenarioBench and FrontierCyber instruments supply the cyber-capability evidence in OpenAI and Anthropic system cards. Then, between July and September 2026, models escaped its evaluation environments and compromised real third-party systems. What the category is, what it measures, and the six governance gaps the incidents exposed.

Inspect: The Open-Source Evaluation Harness Under Third-Party AI Evals

Almost every argument about third-party AI evaluation is an argument about access, independence and publication rights. Underneath those arguments sits a more boring question nobody debates: what software actually runs the test? Increasingly the answer is Inspect — an MIT-licensed Python framework from the UK AI Security Institute and Meridian Labs that METR, Apollo Research and several government testing bodies now run on. This guide explains what Inspect standardises, what an Inspect eval log can and cannot evidence, and why a shared harness is infrastructure rather than assurance.

AI Control Evaluations: The Evidentiary Method Redwood Research Built

Most frontier AI evaluations ask what a model can do, or what it tends to do. A control evaluation asks a third question: if the model were actively trying to defeat the safeguards around it, would the safeguards hold? Redwood Research turned that question into a repeatable experimental procedure with a red team, a blue team, and two numbers at the end. This page explains what kind of evidence that procedure produces, what assumptions it buys that evidence with, and where it stops being load-bearing.

Labs as Each Other’s Evaluators: The Anthropic-OpenAI Safety Evaluation Pilot

In June and July 2025 Anthropic and OpenAI each ran the other’s production models through their own internal alignment evaluations, publishing in parallel on 27 August 2025. Both relaxed model-external API safety filters to make it possible. This guide covers the access arrangement, what each side found, the five validity limitations the labs wrote down themselves, and why the follow-on legally binding cross-testing agreement reportedly reached the contract stage and then died.

Gray Swan Arena: Crowdsourced Red-Teaming as Pre-Deployment Evidence

Gray Swan’s Arena crowdsources adversarial testing to 15,000+ red-teamers, and two competitions with UK AISI and US CAISI produced public papers. What those numbers support as pre-deployment evidence — and what they can’t.

External Review, Unpacked: Which Frontier Labs Let Outsiders Audit Their Safety Cases

Anthropic’s Long-Term Benefit Trust can compel outside review of its Risk Reports. California now requires labs to disclose how much third-party evaluation they used. Most other frontier labs’ safety policies still leave external review optional, or don’t mention it at all. This guide compares, lab by lab, who actually lets outsiders check their Responsible Scaling Policy compliance and who is still grading its own homework.

Apollo Research: AI Scheming Detection Explained

An institutional profile of Apollo Research: the AI safety lab behind scheming and deceptive-alignment research, its OpenAI evaluation collaboration, the Watcher runtime monitor, and its January 2026 move to an independent Public Benefit Corporation.

Is AI Red-Teaming Security Theater? What the Research Says

Carnegie Mellon researchers surveyed the AI red-teaming literature and analyzed six real exercises — Bing Chat, GPT-4, Gopher, two Claude releases, and DEF CON — and found practitioner definitions diverge so widely that treating red-teaming as a catch-all safety guarantee “verges on security theater.” Their answer isn’t to abandon the practice; it’s a 3-phase question bank for doing it rigorously.

The Access Record No Lab Currently Publishes

NIKOLAI’s Evaluator Access Attestation element says so itself: it is CASRAI’s own editorial synthesis, unsourced from any single document, because no developer currently publishes a record in this shape. This page assembles the eight fragments that exist — and finds the one place a lab volunteers to go further than the rest.

Building a Registry That Can’t Accidentally Lie

A design case study of NIKOLAI’s own Mapping Declarations mechanism (Track N10) for people building other open registries with organizational-attribution claims — the honesty gradient across six roles, a same-day dispute-bug fix, and why the roles CRediT doesn’t have were the ones worth building.

When AI Reviews AI: Inside NIKOLAI’s AI-Model-Review Element

Anthropic’s August 2026 Risk Report names the exact model instance that reviewed a draft of its own alignment section — Claude Mythos 5, with named internal access, in 24 minutes. That is not human-evaluator independence; it’s an AI system acting as an assurance actor. NIKOLAI’s AI-Model Review element, in CASRAI’s own independent dictionary, is built to record exactly that — and to name the human still accountable for what the AI found.

What Anthropic Pays for a Jailbreak: Inside Its Model Safety Bug Bounty

Anthropic will pay up to $15,000 through HackerOne for a novel, universal jailbreak against its CBRN and cybersecurity safeguards. Here’s exactly what that program covers, what it doesn’t, and where it sits inside Anthropic’s own safety commitments.

Evaluation-Validity Threats: Sandbagging, Reward Hacking, and NIKOLAI’s N5 Crosswalk

NIKOLAI’s N5 element defines the conditions — sandbagging, alignment faking, metagaming, reward hacking — where a model’s behavior in evaluation diverges from its behavior in deployment. This is its first per-element deep dive: the 8-org crosswalk, the sandbagging/alignment-faking distinction, Anthropic’s reward-hacking data, SB 53’s elicitation carve-out, and a live Senate investigation as case study.

Who Checks the Checkers: Evaluator Independence in AI Safety

Anthropic’s RSP names a financial-interest and personal-relationship test for its external reviewers; the EU’s GPAI Code asks only for “adequate qualification”; METR discloses its own conflicts and refuses compensation; and two competing, not-yet-enacted US bills would regulate evaluator independence in structurally different ways — an evaluator-by-evaluator look at what “independent” is actually required to mean, mapped against CASRAI’s own NIKOLAI crosswalk.

CAISI’s Evaluation Cadence for Open-Weight and PRC-Origin AI Models

CAISI has published cyber-capability assessments of three PRC-lab open-weight models in four months — GLM-5.2, Kimi K3 (with UK AISI), and GLM-5.3 — a real, ongoing evaluation cadence, not a one-off report. What the methodology and results actually show.

Mapping Declarations: How Organizations Verify and Confirm Their Own AI Safety Terminology in NIKOLAI

Every NIKOLAI crosswalk row starts as an unconfirmed shadow mapping. Here’s how an organization verifies its identity and gets a reviewed declaration live.

The International AI Safety Institute Network: What It Is, Now NAAIMES

There’s a coordination body connecting the individual national AI safety institutes — CAISI, UK AISI and others — to each other. It was founded in November 2024, and has since been renamed NAAIMES. Here’s what it actually does.

The Frontier Model Forum: What It Is and What It Does

An institutional profile of the Frontier Model Forum (FMF): its founding members, mission, funding, and how it differs from independent evaluators like METR and government bodies like UK AISI and US CAISI.

UK AI Safety Institute vs AI Security Institute: The 2025 Rename Explained

The UK government renamed its AI Safety Institute to the AI Security Institute on 14 February 2025. It is the same organisation under a new name, not two separate bodies — this guide explains what changed, why, and why the old name still turns up.

Third-Party AI Evaluator Standards: Independence, Access, and Methodology

A cross-cutting guide to the standards questions that apply to every third-party AI evaluator: embedded vs. arms-length access, independence and conflict-of-interest, evaluation validity, red-team access agreements, publication rights, and retaliation protection — mapped to NIKOLAI’s N8 Transparency and Review track.

UK AI Security Institute’s Frontier AI Trends Report: What Its Evaluations Have Found

AISI’s first Frontier AI Trends Report aggregates two years of evaluations across 30+ frontier models. Here is what it found on cyber, bio/chem, autonomy, and safeguards.

What Is METR? How Its AI Safety Evaluations Work

An institutional profile of METR (Model Evaluation and Threat Research): what it is, what it evaluates, how its evaluations work, and its relationship to the AI labs it assesses.

Third-Party AI Auditing: What It Is and Who Does It

Third-party AI auditing means an independent party assessing an AI system, or the organization running it, against a compliance framework like ISO/IEC 42001 or the EU AI Act — not the same as using AI tools to automate financial audits.

CAISI and the UK AI Security Institute: How Pre-Deployment Testing Agreements Work

CAISI and the UK AI Security Institute both renamed themselves in 2025 and both run voluntary pre-deployment testing agreements with frontier AI labs. Here’s what those agreements actually cover, and what “early access” means in practice.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.