Written and maintained by CASRAI Editorial Board
Last updated
What the Frontier AI Trends Report is
On 18 December 2025, the UK AI Security Institute (AISI) published its first Frontier AI Trends Report, aggregating two years of its own government-led model evaluations into a single public analysis. The report draws on internal testing of more than 30 frontier AI systems — mostly general-purpose large language models, with some open-source systems — released between 2022 and October 2025, and evaluated by AISI since it began this work in November 2023.
Where AISI’s identity, mandate, and pre-deployment testing agreements with AI labs are covered in detail in our companion guide, Pre-Deployment Testing: CAISI and UK AISI, this page focuses on what the Trends Report itself found: the specific capability patterns AISI’s evaluators observed across the models they tested, presented with the report’s own figures.
AISI frames the report as the first in an ongoing series. It says the aim is “to improve public understanding about fast-moving AI capabilities and strengthen transparency,” and that future editions will provide iterative, up-to-date visibility into frontier AI development.
Cybersecurity: task completion rates have risen sharply
AISI’s cyber evaluations track how often models can complete offensive-security tasks pitched at different skill levels. The report’s headline figures:
- On apprentice-level cyber tasks, models completed the challenge 9% of the time in late 2023. By the time of the report, that figure had risen to 50%.
- In 2025, AISI tested the first model able to complete cyber tasks written for experts with roughly ten years of professional experience — a difficulty tier no earlier model had cleared.
- Task duration that models can reliably handle has been doubling roughly every eight months.
- Despite this progress, AISI notes models still struggle with cyber-range challenges that require chaining several sequential steps together, rather than solving a single isolated task.
Chemistry and biology: models have overtaken PhD-level baselines
AISI’s science evaluations compare model performance against human expert baselines on open-ended biology and chemistry questions, and against protocol-generation and troubleshooting tasks used in wet-lab contexts.
- AISI first tested a model surpassing the average score of biology PhD holders (40–50%) in 2024. By the report’s writing, frontier models “far surpassed” that PhD-level baseline, exceeding it by up to 60% on some open-ended question sets, with chemistry performance catching up quickly.
- Models first produced experiment protocols judged accurate by AISI’s reviewers in late 2024.
- On troubleshooting tasks — diagnosing what went wrong in a failed lab procedure — models scored up to 90% better than the human experts AISI compared them against.
- In controlled evaluations, non-specialists assisted by a frontier model were 4.7 times more likely to succeed at generating an accurate viral-recovery protocol than those relying on unaided internet research.
- By mid-2025, models first outperformed human experts on multimodal science tasks (those combining text with images, such as reading lab equipment or gel electrophoresis results).
Autonomous task completion and self-replication
AISI also evaluates how far models can operate independently on multi-step tasks and, more narrowly, on the specific sub-tasks involved in replicating themselves without human help.
- On tasks that would take a skilled human roughly an hour, models achieved under 5% success in late 2023. By mid-2025, success on comparable tasks exceeded 40%.
- Using AISI’s RepliBench benchmark for self-replication sub-tasks, success rates rose from around 5% in 2023 to 60% in 2025 on average, and two frontier models exceeded 60% overall.
- AISI reports models are strongest at the early stages of self-replication (obtaining compute and money) and weaker at later stages of the process.
- The report separately notes that AI agents with execution-capable access to financial infrastructure — what AISI categorises as Levels 4–5 autonomy — are increasingly common among newly released systems.
The open-source gap is narrowing
AISI compared how long it takes an open-source model to match the capability of the frontier closed model that preceded it. Estimates vary by methodology: roughly four months per Artificial Analysis’s tracking, or eight months per METR’s. AISI cites both, giving a four-to-eight-month range for how far behind frontier open-source models now trail closed ones — a gap the report says has been narrowing over the two years it covers.
Safeguards: jailbreaks persist, and protection varies widely between models
AISI’s safeguard testing sits alongside its capability testing. Two findings stand out:
- AISI says it found a universal jailbreak — a way to bypass safety training entirely — in every single system it tested.
- The effort required to find a jailbreak for biological-misuse content varied enormously between models: one model released six months after its predecessor required roughly 40 times more expert effort to jailbreak in this way.
- Across the models it tested, AISI found only a weak statistical relationship between how capable a model is and how well-defended it is (R² = 0.097) — meaning stronger capability does not reliably predict stronger safeguards.
Societal use patterns AISI also tracked
Alongside capability and safeguard testing, the report includes survey-based findings on how the UK public is actually using these systems:
- In a survey of 2,028 UK participants, 33% said they had used an AI model for emotional or companionship purposes in the past year; 8% did so weekly and 4% daily.
- 32% of chatbot users said they had researched election-related topics using a chatbot ahead of the UK’s 2024 general election.
Methodology notes
AISI states that each task within each evaluation was repeated ten times, and that the report presents aggregated results across its internal evaluation suite rather than a single benchmark run. Because it draws on AISI’s own internal testing infrastructure, the same tasks are not independently reproducible in the way a public leaderboard benchmark would be — the report itself is the evidence, and AISI has said it intends to publish further editions on an iterative basis.
Why this matters for organisations tracking frontier AI risk
The Trends Report is a rare public dataset from a government evaluator with direct pre-deployment access to frontier labs, rather than a benchmark run by a vendor or an academic group with only API access. For organisations mapping which capability categories to monitor — cyber offense, biological and chemical uplift, autonomous operation, and safeguard robustness — the report gives some of the only publicly available trend lines showing how these categories have moved over a defined two-year period, rather than a single snapshot.
The underlying question of who UK AISI is, how its pre-deployment testing agreements with labs work, and how it relates to the US CAISI, is covered separately in our guide on pre-deployment testing arrangements between CAISI and UK AISI. For organisations building governance processes around evaluation evidence like this, NIKOLAI provides a structured vocabulary for describing and mapping AI risk-management elements referenced in frameworks and evaluation regimes of this kind.
FAQ
What is the Frontier AI Trends Report?
It is the UK AI Security Institute’s first public report aggregating two years of its own frontier-model evaluations, published 18 December 2025. It covers more than 30 systems tested since November 2023 across cybersecurity, chemistry and biology, autonomous task completion, and safeguard robustness.
Is this the same as UK AISI’s pre-deployment testing agreements with labs?
No. The pre-deployment testing agreements are bilateral arrangements between AISI and individual labs for pre-release model access. The Trends Report is a separate publication that aggregates findings from AISI’s evaluation work, some of which draws on models tested under those agreements and some on models tested after public release.
How much have cyber capabilities changed, according to the report?
AISI reports that apprentice-level cyber task completion rose from 9% in late 2023 to 50% by the report’s writing, and that it tested the first model able to complete expert-level cyber tasks in 2025.
Have AI models really surpassed PhD-level biology experts?
According to AISI’s own evaluations, yes on the specific open-ended question sets and troubleshooting tasks it tests against a PhD-holder baseline. AISI first observed a model surpassing that baseline in 2024, and by the report’s writing said frontier models exceeded it by up to 60% on some question sets.
Does AISI say frontier models are unsafe?
The report doesn’t make a single up-or-down safety judgment. It documents rising capability alongside persistent jailbreak vulnerabilities across every system tested, and a weak correlation between capability and safeguard strength — framed as evidence for policymakers and the public rather than a pass/fail verdict.
Will AISI publish more Trends Reports?
AISI has said this is the first in an intended series, with future editions providing iterative, up-to-date public visibility into frontier AI development.







