Written and maintained by CASRAI Editorial Board
Last updated
Last verified against primary sources: 25 September 2026. Debates about third-party AI evaluation almost always run on the governance layer — who gets model access, on what terms, with what right to publish. CASRAI covers that layer in its guides to third-party evaluator standards and methodology and evaluator independence. Underneath it sits a question that gets almost no attention and quietly determines whether any of it is reproducible: what software actually runs the test?
For a growing share of frontier-model safety testing, the answer is Inspect. Its documentation describes it plainly: “Inspect is a framework for frontier AI evaluations developed by the UK AI Security Institute and Meridian Labs.” It is a Python package, MIT-licensed, requiring Python 3.10 or later; as of 22 September 2026 the published version on PyPI was 0.3.268, the 246th release since the project was first open-sourced in May 2024. Its one-line summary on PyPI is unglamorous and accurate: “Framework for large language model evaluations.”
This page is about what that framework standardises, what artefacts it produces that an outside reviewer can actually check, and — the part that matters most for governance — the specific validity problems a shared harness does not fix. NIKOLAI, CASRAI’s own independent frontier-AI-safety dictionary, has two N5 elements (evaluation run and elicitation method) that map onto the exact things an Inspect log does and does not capture; more on that below.
Who builds it, and why the attribution changed
Inspect was created at the UK AI Safety Institute — renamed the AI Security Institute in February 2025, a rebrand CASRAI covers in UK AI Safety Institute vs. AI Security Institute — and released publicly in May 2024. Its principal author, JJ Allaire, joined AISI as a staff engineer to lead the work, and in 2025 co-founded Meridian Labs to build open-source tooling for AI research and evaluation. The framework is now credited jointly to AISI and Meridian Labs, while the repository still lives under the UKGovernmentBEIS GitHub organisation, a fossil of the department AISI was originally established within.
That joint attribution is not a cosmetic detail for anyone assessing governance risk. A testing harness maintained solely inside a national security institute is a different dependency from one maintained by a government body and an independent lab together, under a permissive licence, in public. The MIT licence in particular means nothing in the tooling layer is conditioned on a relationship with the UK government: a commercial evaluator, an academic group or a competing state institute can fork it without asking.
The anatomy of an Inspect evaluation
Inspect’s conceptual model is deliberately small. An evaluation is a task, and a task is three things bolted together:
| Component | What it is | Why governance cares |
|---|---|---|
| Dataset | Labelled samples, “typically a table with input and target columns” |
Defines what is being asked. The contamination and canary-string question lives here. |
| Solver | The thing that “produce[s] an answer for each sample” — from a bare model call up to a full agent | This is the elicitation method. Swap the solver and the same dataset yields a different capability estimate. |
| Scorer | Evaluates the output “using text comparisons, model grading, or other custom schemes” | Decides what counts as success. Scorer choice is a substantive judgement dressed as a technicality. |
The built-in scorers are worth naming, because the list is short enough to reason about: includes(), match(), pattern(), answer(), exact(), f1(), choice(), math(), perplexity(), target_perplexity(), and the two model-graded scorers model_graded_qa() and model_graded_fact(). Most report accuracy and stderr; exact() and f1() report mean and stderr. The presence of stderr as a default metric is a small but real piece of discipline — a harness that reports uncertainty by default makes it slightly harder to publish a bare percentage with no error bar.
The model-graded pair deserves a governance flag. model_graded_qa() is recommended “for open-ended answers” and model_graded_fact() for outputs “too complex to assess with match() or pattern()”. Both mean a language model is deciding whether another language model passed. That is often the only tractable option, but it converts the grader into an undeclared dependency of the result, and a commissioner reading a score should know which grader model produced it.
Agents, tools and sandboxing: the part that makes agentic evals portable
Simple question-answering benchmarks barely need a framework. The reason Inspect spread is that agentic evaluation — the kind that actually probes autonomous capability — is an infrastructure problem rather than a prompting problem, and everyone was solving it separately.
Inspect ships a react() agent implementing a reason-act-observe loop, plus deep agents that “combine planning, memory, and tool use”. It provides built-in bash(), Python, text-editing, web search, web browsing, computer-use and todo_write() tools, and accepts custom and MCP tools alongside them. Crucially, it abstracts the sandbox: the documentation lists support for Docker, Kubernetes, Modal, Proxmox and Vagrant. It talks to over 20 model providers, plus local inference.
The governance consequence is portability. When a dangerous-capability evaluation is written against Inspect, a different organisation with different model access and different compute can re-run the same task definition against the same sandbox semantics. That is the precondition for anything resembling replication in this field, and it did not previously exist. It is also why METR migrated off its in-house Vivaria platform onto Inspect, and why Apollo Research deprecated its own internal evaluation framework after testing it — both changes AISI reported in its June 2026 engineering playbook post.
The eval log is the actual deliverable
For assurance purposes, the most important object Inspect produces is not the score. It is the log.
Every run emits an EvalLog with a defined structure. The named top-level fields include version (the file-format version, currently 2), status (started, success or error), eval (task, model and creation timestamp), plan (the solvers and generation configuration actually used), results (aggregate scores from the scorer metrics), stats (token usage), samples (every individual sample with its input, output, target and score), reductions (how per-sample values were reduced across multiple epochs), error, and tags/metadata. There is also a log_updates field holding an append-only edit history with provenance tracking — so a post-hoc annotation is visible as an annotation rather than silently replacing what was recorded at run time.
Logs write either as .eval, a binary format roughly one-eighth the size of JSON with incremental sample access, or as .json for human readability; both work with the same log API and can be mixed in one directory. A browser-based Inspect View tool renders them for inspection, and there is a VS Code extension for authoring and debugging.
Read that field list through a records-management lens and what Inspect has really shipped is a file format for evaluation evidence. The model identifier, the generation config, the solver chain, every sample and every score are in one artefact that travels. An evaluator who hands a commissioner a PDF summary is asking to be trusted; an evaluator who hands over the .eval files is handing over something checkable.
Inspect also supports running sets of evaluations via eval_set() / inspect eval-set, which retries failed evaluations under a configurable strategy (default 10 attempts), re-uses samples from failed tasks so work is not repeated, cleans up logs from failed runs, and uses the log directory itself as “a durable record of which tasks are completed” so a long run can be resumed. One documented sharp edge: sample deduplication relies on explicit id fields or stable ordering, and the docs warn it “will not work correctly if your dataset is shuffled”.
Inspect Evals: the shared register of benchmarks
Inspect is the harness; Inspect Evals is the library of evaluations built on it. Announced on 13 November 2024, it began as a collection of community-contributed benchmark implementations so that, in AISI’s words, “these benchmarks can now be run against any model with a single command.”
As of September 2026 the published register listed 171 evaluations — 129 internal and 42 external, spanning cybersecurity (Cybench, AgentDojo, CyberGym), coding and software engineering (HumanEval, MBPP, SWE-bench, Terminal-Bench), scientific reasoning (LAB-Bench, SciCode, HealthBench), knowledge and reasoning (MMLU, GPQA, ARC, GAIA), safety and safeguards (StrongREJECT, MASK, XSTest, FORTRESS), multimodal tasks, mathematics, and bias and fairness suites (BBQ, BOLD, StereoSet). The repository is MIT-licensed and is now maintained by Generality Labs, a London-based nonprofit, with founding support credited to the UK AI Security Institute, Arcadia Impact and the Vector Institute. Direct pull requests are restricted to pre-approved contributors.
Note what this does and does not give you. A shared register makes benchmark implementations comparable, which removes a genuine source of spurious disagreement — two labs reporting different SWE-bench numbers because they implemented it differently. It does nothing about whether the benchmark measures what its name suggests.
Where a harness becomes a standard: the Autonomous Systems Evaluation Standard
The clearest example of Inspect being used as more than a convenience is UK AISI’s Autonomous Systems Evaluation Standard, published 31 October 2024. It sets minimum requirements for evaluations submitted into AISI’s autonomous-systems suite, and its first requirement is categorical: “All evaluations must be built using Inspect.”
The rest of the standard is the interesting part, because it is quality control that a harness alone cannot impose:
- Repository structure. Submissions must follow a cookiecutter template and pass automated checks including type annotations and formatting.
- Documentation. Three required files — a
README.mdthat must contain a canary string so the evaluation can be detected if it ends up in training data, aCONTRIBUTING.mdfor maintenance, and a structuredMETADATA.json. - Automated scoring. “Scoring for your evaluation should be fully automated, requiring no manual grading steps.”
- Evidence of QA. Developers must supply Inspect log files showing at least one completed sample, plus manual verification against frontier models.
- Safety of the task itself. “You should not encourage the agent to do anything unsafe, illegal, or unethical in the real world.”
The canary-string requirement is the one governance readers should register. It is an explicit admission that benchmark contamination is a live, expected failure mode, and it builds a detection affordance into the submission format rather than leaving it to hope. The QA-log requirement is the second: the standard treats an Inspect log as the acceptable evidence that an evaluation actually runs. That is a small, concrete instance of an evaluation artefact being used as an assurance artefact.
What a shared harness does not solve
This is the section that matters if you are drawing governance conclusions from the fact that an evaluation “was run in Inspect”.
1. Elicitation is not standardised, and it dominates results. The solver is a free parameter. The same dataset, the same scorer and the same model produce materially different numbers depending on the scaffold, the tool set, the number of attempts and the prompt. A harness makes the elicitation method recordable; it does not make it comparable. NIKOLAI’s elicitation method element exists precisely because this parameter is usually reported informally or not at all, and because a capability number means something different depending on whether it is offered as a lower bound or a ceiling.
2. Outcome scores conceal more than they reveal. The May 2026 paper Log analysis is necessary for credible evaluation of AI agents — whose author list includes JJ Allaire alongside Sayash Kapoor, Arvind Narayanan, Marius Hobbhahn, Jacob Steinhardt, Cozmin Ududec, Magda Dubois, Conrad Stosz and others — argues that outcome-only benchmarking produces inflated scores from shortcuts, poor predictions of real-world performance, and concealed dangerous behaviours. On tau-Bench Airline the authors found headline metrics understated capability by roughly 50% while hiding deployment problems invisible to the score. That an Inspect co-author is making this argument is the point: the framework’s own people are saying the log, not the number, is the evidence. It sits squarely alongside the critique CASRAI summarises in AI safety is not a model property.
3. Scorer validity is out of scope. Nothing in a harness tells you whether model_graded_qa() agreed with a human expert on the outputs that mattered. CASRAI’s guide to NIKOLAI N5 evaluation-validity threats works through this class of problem in detail.
4. A common harness is not independence. An internal safety team and an external evaluator can run byte-identical Inspect tasks and still differ on everything that matters — what was tested, when, against which checkpoint, and what happened to an unflattering finding. Tooling convergence can even mask a governance gap, because two parties agreeing on software reads as agreement.
5. Concentration risk is real. If METR, Apollo Research, UK AISI and US CAISI converge on one harness, a bug or a design assumption in that harness propagates across the ecosystem at once. The MIT licence and open development mitigate this; they do not remove it.
Where NIKOLAI fits
NIKOLAI is CASRAI’s own independent frontier-AI-safety dictionary. It is not endorsed by UK AISI, Meridian Labs or any of the organisations named on this page, and no crosswalk row here reflects an agreement with them — any mapping between NIKOLAI terms and Inspect’s vocabulary is a shadow mapping until an organisation files a Mapping Declaration.
Read as a shadow mapping, though, the fit is unusually tight, because Inspect independently converged on the same distinctions NIKOLAI’s N5 track draws:
| NIKOLAI element (N5 · Evidence and evaluations) | What NIKOLAI records | Inspect analogue (shadow mapping) |
|---|---|---|
| Evaluation run | “A single execution, or a declared batch of executions, of an evaluation against a specified model checkpoint and configuration, recorded with attempt count, scoring rule, and date.” | An EvalLog: eval (task, model, timestamp), plan (config), reductions (epochs), results (scorer metrics) |
| Elicitation method | The techniques and conditions used to draw out maximum capability, recorded as a property of the run, with the declared interpretation (lower bound vs. ceiling) | The plan field’s solver chain, agent and tool set — recorded, but with no declared interpretation attached |
| Evaluator independence | Conflict-of-interest terms governing third-party review | No analogue. Out of scope for any harness. |
The gap in row two is the useful observation. Inspect records what was done with high fidelity. NIKOLAI’s element asks for something a log cannot supply on its own: the evaluator’s declared reading of the result. A 40% score produced by a deliberately weak scaffold and a 40% score produced by a best-effort elicitation attempt are the same number and completely different evidence, and only the evaluator can say which they were running.
The research-administration angle
There is a real one here, and it is narrower than the topic’s prominence suggests.
Academic groups now run frontier-model evaluations, and the Vector Institute’s role in founding Inspect Evals is one visible example. When a university group takes on that work, it lands on research administration in three specific places. First, research computing: Inspect’s sandbox requirements are Docker or Kubernetes with untrusted, model-generated code executing inside them, which is a materially different request to a central IT or HPC team than a batch compute allocation, and it usually needs an explicit conversation about network egress and isolation. Second, research security and export control: evaluations of dual-use capability — the cyber-offence and bio suites in particular — involve generating and storing artefacts a research security office should see before, not after, a publication decision. Third, research data management, where the fit is closest: an .eval log is a structured, versioned, provenance-bearing record of a computational experiment, and the question of where it is deposited, for how long, and whether it accompanies the publication is the same question an RDM office already answers for every other computational result.
What does not apply: model evaluation is not human-subjects research, and routing an Inspect-based capability evaluation to an IRB is a category error unless human participants are actually involved in scoring or red-teaming.
Frequently asked questions
Is Inspect free to use commercially? Yes. Both inspect_ai and the inspect_evals register are MIT-licensed, which permits commercial use, modification and redistribution with attribution and no reciprocal obligations.
Does using Inspect make an evaluation independent? No. Inspect is a test harness. Independence is a property of the relationship between the evaluator and the developer — access terms, funding, conflict-of-interest rules and publication rights — and none of those are software features. See evaluator independence in the AI safety ecosystem.
Do regulators require Inspect? No general-purpose AI regulation names it. The closest thing to a mandate is UK AISI’s Autonomous Systems Evaluation Standard, which requires Inspect for evaluations submitted into AISI’s own autonomous-systems suite — a procurement-style condition on a specific submission channel, not a legal requirement on developers.
What is the difference between Inspect and Inspect Evals? Inspect is the framework that runs evaluations. Inspect Evals is a separate MIT-licensed repository of benchmark implementations built on it, maintained by Generality Labs, listing 171 evaluations as of September 2026.
Can a commissioner verify an evaluation from the log alone? Partly. The log shows which model and configuration ran, every sample, and every score, which is far more than a summary report. It does not show why that elicitation method was chosen, what else was tried and not reported, or whether the benchmark was in the training data — which is why the AS standard adds a canary string and why the log-analysis literature treats logs as necessary but not sufficient.
Who maintains Inspect now? It is developed jointly by the UK AI Security Institute and Meridian Labs, the open-source lab JJ Allaire co-founded in 2025, with the repository hosted under the UKGovernmentBEIS GitHub organisation.
Primary sources
- Inspect documentation — inspect.aisi.org.uk (framework description, components, agents, tools, sandboxing, eval logs, scorers, eval sets)
- Inspect source repository — github.com/UKGovernmentBEIS/inspect_ai (MIT licence)
- PyPI: inspect-ai — version 0.3.268, released 22 September 2026; Python ≥3.10; 246 releases
- UK AISI, Open-sourcing our testing framework Inspect (2024)
- UK AISI, Announcing Inspect Evals, 13 November 2024
- UK AISI, Releasing AISI’s Engineering Playbook, 18 June 2026 (METR and Apollo Research migrations)
- UK AISI, Autonomous Systems Evaluation Standard, 31 October 2024
- Inspect Evals register (171 evaluations; maintained by Generality Labs)
- Kirgis, Kapoor, Rabanser, Nadgir, Ududec, Dubois, Allaire, Stosz, Hobbhahn, Steinhardt and Narayanan, Log analysis is necessary for credible evaluation of AI agents, May 2026








