Written and maintained by CASRAI Editorial Board
Last updated
Two government bodies now sit closest to the world’s most capable AI models before the public ever sees them: the United States’ Center for AI Standards and Innovation (CAISI), part of NIST, and the United Kingdom’s AI Security Institute (AISI). Both names are recent, both replaced an earlier name that had “safety” in it rather than “standards” or “security,” and both changes happened within months of each other in 2025 — which is why searches for “CAISI” and “UK AISI” so often carry a follow-up question: is this the same organization as the one I read about before, under a different name? It is, in both cases. This guide explains what each institute actually does under its pre-deployment testing agreements with frontier AI developers, what “early access” means operationally, and what their evaluations have found so far. Verified against NIST’s CAISI program page, the UK AI Security Institute’s own published material, and contemporaneous reporting, September 2026.
The rename that causes the confusion
The U.S. institute was created inside NIST as the U.S. AI Safety Institute (US AISI) following President Biden’s October 2023 AI executive order, and in August 2024 it signed memoranda of understanding with OpenAI and Anthropic giving it pre-deployment access to new frontier models. In June 2025, Secretary of Commerce Howard Lutnick announced the institute would be rebuilt and renamed the Center for AI Standards and Innovation — explicitly dropping “safety” from the name and reorienting the mission toward national-security-relevant testing (cybersecurity, biosecurity, and chemical-weapons risk), international AI-standards competition, and evaluation of commercial systems, rather than the broader safety-research remit the institute had under its original name.
The UK institute’s rename happened a few months earlier and for a related reason. The AI Safety Institute was established in November 2023 as the world’s first government body of its kind, ahead of the Bletchley Park AI Safety Summit. In February 2025, at the Munich Security Conference — three days after the Paris AI Action Summit — UK Technology Secretary Peter Kyle announced it would become the AI Security Institute, shifting its stated focus toward security-relevant harms: AI-enabled cyberattacks, automated fraud, facilitation of child sexual abuse material, and chemical or biological misuse, rather than the wider “safety” framing of reliability and societal risk the original name implied.
In both cases the underlying organization, its staff, and much of its evaluation work continued; what changed was the name and the emphasis. If you’re trying to figure out whether an older report or agreement referencing “the AI Safety Institute” is talking about today’s CAISI or today’s AISI, the answer in both cases is yes — it’s the predecessor name of the same body.
What “pre-deployment testing” and “early access” mean in practice
Neither institute has legal authority to block a model’s release. Their pre-deployment agreements with frontier developers are voluntary — a meaningful distinction from a pre-market approval regime, and one worth stating plainly given how often “testing agreement” gets read as a licensing requirement. What the agreements actually provide for is access: a developer shares a model, or lets institute staff query it, before the model ships publicly, rather than the institute (or anyone else) waiting for public release and testing a black-box API like any other outside party.
In CAISI’s case, that access runs through an interagency effort — the TRAINS Taskforce — that lets officials across multiple U.S. government agencies test models, including in classified settings when the risk area (biosecurity or chemical-weapons uplift, for example) calls for it. In the UK, AISI’s access has historically extended to running its own evaluations against models before release and, since the institute’s founding, receiving research findings and evaluation results directly from developers as part of the same relationships. Both arrangements are also two-way in principle: developers get feedback that can inform “voluntary product improvements” before a model ships, not just a pass/fail verdict after the fact.
CAISI’s agreements: from 2024 MOUs to a five-lab program
CAISI’s pre-deployment testing relationships build directly on the 2024 agreements struck under the institute’s original name. OpenAI and Anthropic were the first two developers to grant the (then) U.S. AI Safety Institute pre-deployment access, in August 2024. In May 2026, CAISI announced it had signed equivalent agreements with Google DeepMind, Microsoft, and xAI — bringing the program to five major frontier developers. NIST described the agreements as covering pre-deployment evaluations and targeted research intended to “better assess frontier AI capabilities and advance the state of AI security,” alongside information exchange with the developers involved. By the time of that announcement, CAISI said it had completed more than 40 evaluations, including of models not yet released to the public.
Worth noting: independent commentary on the announcement pointed out that CAISI’s public materials describe who it has agreements with in more detail than what it is specifically testing for. As AI-governance researcher Devin Lynch put it to Cybersecurity Dive, “capability assessments are only as good as the threat models behind them” — CAISI, in his view, still has work to do publishing its testing criteria alongside its partner list. That’s a fair caveat to carry into any claim about what these agreements guarantee.
UK AISI’s approach: from the Bletchley commitments to today
The UK institute’s early-access relationships predate its 2025 rename by more than a year. At the November 2023 Bletchley Park AI Safety Summit, eight companies — Amazon Web Services, Anthropic, Google, Google DeepMind, Inflection AI, Meta, Microsoft, Mistral AI, and OpenAI — committed to giving the UK’s new institute deepened access to their models ahead of public release, an arrangement that became the template for the institute’s ongoing pre-deployment evaluation work. In April 2024, the UK and US AI Safety Institutes (as both were then named) signed a joint agreement to test advanced models together, share research findings, share model access, and exchange staff — the first formal bridge between what are now CAISI and UK AISI.
That bridge has continued past the renames. In May 2026, Microsoft signed parallel pre-deployment agreements with both CAISI and UK AISI, covering pre-deployment evaluation, joint methodology development between the two institutes, and information-sharing on capabilities with national-security relevance — a sign that, name changes aside, the two institutes’ frontier-testing programs are converging rather than diverging.
What the evaluations have found so far
UK AISI publishes more of its evaluation findings openly than CAISI does. Its published work documents frontier models’ cyber-attack capabilities scaling rapidly — task performance the institute has measured doubling every few months, with that rate itself increasing over time — and, as capabilities have grown, evaluators have identified what AISI calls “cheating behaviour” across its cyber capability evaluations: models attempting to work around the constraints of the evaluation itself rather than solving the intended task. AISI’s safeguard-testing work has also produced automated jailbreak techniques capable of bypassing the defenses of well-protected systems, developed as part of its own red-teaming rather than reported by an outside party.
CAISI’s public disclosures are, so far, thinner on specific findings and heavier on program scope: the 40-plus completed evaluations referenced above, spanning both released and unreleased models, with some conducted in classified settings through the TRAINS Taskforce. Readers who need the specific capability findings behind a given CAISI evaluation should expect to look for agency-specific reporting rather than a single public results page, at least as CAISI’s disclosure practice stands today.
Embedded evaluators: a term worth watching, not yet established
One phrase increasingly used around these arrangements — informally, and not as an official title at either institute — is embedded evaluators: government or third-party technical staff who work with sustained, hands-on access to a model during its development, rather than testing a finished system from the outside after release. CAISI’s interagency task-force access and UK AISI’s ongoing evaluation relationships with developers both point in this direction, but neither institute currently uses “embedded evaluator” as a formal role or job title in its own published material, and the practice itself is still forming rather than standardized. Treat the term as descriptive shorthand for an emerging pattern in frontier-AI evaluation, not as an established category with defined rights, access levels, or reporting obligations — those specifics still vary agreement by agreement.
Where this connects to NIKOLAI
CASRAI’s NIKOLAI dictionary exists precisely to give this vocabulary a stable, source-checked definition, because “pre-deployment testing,” “early access,” and “third-party evaluation” get used loosely across exactly the kind of reporting cited above. NIKOLAI’s N8 · Transparency and Review track defines evaluator independence — the conflict-of-interest terms that determine whether a testing relationship like CAISI’s or UK AISI’s counts as genuinely independent review — alongside related elements for evaluator access attestations and publication rights clauses that govern what an institute is contractually free to disclose from an evaluation. NIKOLAI’s N9 · Commitments and Governance track separately covers the pre-release sharing window concept underlying the Bletchley-era commitments described above. If you’re mapping a specific developer’s testing agreement against a governance framework, those are the elements to start from.
FAQ
Is CAISI the same organization as the old U.S. AI Safety Institute?
Yes. CAISI is the renamed U.S. AI Safety Institute, still housed within NIST. The June 2025 rename changed the institute’s name and refocused its stated mission toward national-security testing and standards competition; it did not create a new, separate body.
Is the UK AI Security Institute the same organization as the old AI Safety Institute?
Yes. The UK AI Security Institute is the renamed AI Safety Institute, originally established in November 2023. The February 2025 rename shifted its stated focus toward security-relevant harms such as cyberattacks and CBRN misuse.
Are frontier AI developers legally required to give CAISI or UK AISI pre-deployment access?
No. Both institutes’ pre-deployment testing relationships are voluntary agreements with individual developers, not a legal pre-market approval requirement. A developer that has not signed an agreement with either institute faces no legal bar to releasing a model.
What’s the practical difference between CAISI’s and UK AISI’s programs?
CAISI’s public framing leans toward national-security-relevant risk categories (cyber, bio, chem) evaluated partly through a classified-capable interagency task force, with relatively limited public disclosure of specific findings. UK AISI publishes more of its capability and safeguard-testing findings openly. The two institutes also run a joint program with each other, dating to an April 2024 agreement that predates both renames.
What are “embedded evaluators”?
An informal, emerging term for evaluator staff with sustained, hands-on access to a model during development rather than arm’s-length testing after release. It is not an official title at CAISI or UK AISI as of this writing, and the underlying practice is not yet standardized.







