Skip to main content
v2026.11,858 entries · CC-BY 4.0

What Is a Frontier AI Model?

“Frontier AI model” is a capability class, not a specific product or lab. This guide walks through the three definitional approaches in active use — compute thresholds, capability thresholds, and relative state-of-the-art — and why the definition a framework uses is usually what triggers its obligations.

Written and maintained by CASRAI Editorial Board

Last updated

A “frontier AI model” is not a specific product, and it is not a specific company’s model. It is a capability class: a model at or near the current edge of what AI systems can do. What varies — and what actually matters if you build, deploy, or regulate one — is which definition applies, because three different definitional approaches are in active use across current law and industry practice, and they draw the line in different places for different reasons.

The short answer

Depending on who is asking, a model becomes “frontier” when it crosses one of three kinds of line:

  • A compute threshold — a fixed amount of training compute, measured in floating-point operations (FLOPs), written into statute. This is the approach California’s SB 53, New York’s RAISE Act, and the EU AI Act all use.
  • A capability threshold — a demonstrated dangerous capability (e.g. uplift toward CBRN weapons development, or the ability to automate AI research itself), evaluated directly rather than inferred from compute spend. This is how Anthropic’s Responsible Scaling Policy defines when a model needs stronger safeguards.
  • A relative, state-of-the-art threshold — whether a model outperforms everything that has already been widely deployed for some period of time. This is the Frontier Model Forum’s framing.

None of these three approaches is “the” definition. They coexist, they were written for different purposes, and a single model can be frontier under one and not another at the same time.

A capability class, not a lab or a product

It’s worth being precise about what the term is not. “Frontier AI model” doesn’t name a fixed list of specific systems from specific labs — it names a moving category. A model that counted as frontier in one year can fall out of the category later, not because it changed, but because the frontier moved past it. Every compute-threshold and capability-threshold definition below is written to track that movement: the EU AI Act’s Commission can revise its FLOPs threshold by delegated act as training costs fall, and Anthropic’s Responsible Scaling Policy is explicitly versioned and updated as capabilities change. Treat any claim that a specific named model “is” or “isn’t” frontier as time-stamped, not permanent.

Three ways the term gets defined

1. Compute thresholds: a bright line measured in FLOPs

The dominant approach in current US and EU law is to pick a fixed amount of training compute and treat anything above it as frontier, regardless of what the model actually does. The appeal is administrability: compute spend is something a developer can (in principle) measure and disclose, where “dangerous capability” requires an evaluation.

  • California SB 53 (the Transparency in Frontier Artificial Intelligence Act, signed into law September 29, 2025) defines a “frontier model” as a foundation model trained using more than 1026 integer or floating-point operations — counting the original training run plus any subsequent fine-tuning, reinforcement learning, or other material modification. A “frontier developer” is whoever trained it or intends to; a “large frontier developer” is a frontier developer whose group had annual gross revenue over $500 million in the preceding year, and it’s the “large” tier that carries the heavier obligations.
  • New York’s RAISE Act uses the identical 1026 operations threshold and the same $500 million revenue line for “large frontier developer,” deliberately aligned with California’s after New York amended the bill to match. See the RAISE Act guide for how New York’s obligations differ from California’s despite the shared definition.
  • The EU AI Act sets its bar lower and uses it differently. A general-purpose AI model trained with more than 1023 FLOPs is presumptively “general-purpose” under Article 3(63) in the first place; a model trained with more than 1025 FLOPs is presumed under Article 51(1)(a) to pose “systemic risk,” which is the EU’s closest analogue to “frontier.” The European Commission can also designate a model as systemic-risk on other grounds — capability evaluations, reach, or tool access under Annex XIII — independent of the FLOPs count, and can revise the thresholds themselves over time.

2. Capability thresholds: what the model can actually do

Compute is a proxy. Some frameworks skip the proxy and evaluate the thing regulators actually care about directly. Anthropic’s Responsible Scaling Policy defines its AI Safety Levels (ASL) this way: instead of a FLOPs number, it specifies Capability Thresholds — concrete, evaluated capabilities such as the ability to meaningfully uplift a moderately resourced state CBRN weapons program, or the ability to fully automate entry-level AI research work, or to compress roughly two years of 2018–2024-pace AI progress into a single year. Crossing one of those evaluated thresholds, not a compute number, is what requires moving from the ASL-2 baseline to the ASL-3 Security and Deployment Standards. See the Responsible Scaling Policy guide for how the ASL levels and their evaluations work in practice.

This matters because a capability-threshold definition and a compute-threshold definition don’t have to move together. A model can cross SB 53’s 1026-operation line without ever being evaluated as crossing an ASL-3 capability threshold, and in principle a smaller, more efficiently trained model could demonstrate a dangerous capability without approaching 1026 operations at all.

3. Relative thresholds: state-of-the-art compared to what’s already out

The Frontier Model Forum — the industry body founded by Anthropic, Google, Microsoft, and OpenAI — defines a frontier AI model as a general-purpose model that outperforms, on conventional benchmarks or high-risk capability assessments, every other model that has been widely deployed for at least twelve months. “Frontier AI” more broadly is simply whatever constitutes the state of the art at a given moment — a collection that shifts as the field moves, by construction rather than by any fixed number. This definition has no compute or capability floor at all; it’s purely comparative.

Why the definition is the trigger: what crossing the line actually requires

None of this is academic, because the definition is usually the whole mechanism: crossing the threshold is what switches on a framework’s obligations. SB 53 is the clearest illustration, because its obligations scale directly with which definitional tier a developer falls into.

Every frontier developer under SB 53 must report “critical safety incidents” to California’s Office of Emergency Services within 15 days of discovery (24 hours if there’s imminent risk of death or serious injury), and must publish a transparency report before deploying a frontier model, covering contact information, release date, supported languages, output modalities, intended uses, and usage restrictions.

Large frontier developers — the ones over the $500 million revenue line — additionally have to write, publish, and annually update a documented frontier AI framework covering how they define catastrophic risk thresholds, apply mitigations, use third-party evaluators, and handle cybersecurity and incident response; publish material changes to that framework within 30 days with a justification; expand their transparency reports to include catastrophic risk assessment summaries and third-party involvement; send quarterly risk-assessment summaries to the Office of Emergency Services; and maintain an anonymous internal whistleblower channel with monthly status updates.

Under the EU AI Act, crossing into “systemic risk” territory brings its own separate set of obligations (notification to the AI Office, model evaluation, incident reporting, cybersecurity requirements) on top of the baseline duties every general-purpose AI model provider already has. The mechanism is the same shape in both jurisdictions even though the numeric thresholds differ: the definition isn’t just descriptive, it’s the switch.

Why the same model can be “frontier” under one framework and not another

Because these three approaches measure different things, a developer training a genuinely capable model can find itself in scope of a compute-threshold law, out of scope of a capability-threshold framework, and described as “frontier” by an industry body’s relative definition — all at once, all correctly, none of them contradicting the others. There is no single authority that reconciles these into one number. Anyone reading a claim that “X is a frontier model” needs to ask: frontier under whose definition, and as of when?

Where NIKOLAI fits

NIKOLAI is CASRAI’s open, versioned dictionary of elements for frontier-AI safety documentation. It exists precisely because of the fragmentation described above: its N1 track standardizes how different frameworks describe actors, models, and scope, and its N3 track standardizes how they describe thresholds and checkpoints — including the kind of compute and capability thresholds covered here — so that a “frontier model” claim made under SB 53, the EU AI Act, or a lab’s own Responsible Scaling Policy can be crosswalked against the others using a shared vocabulary. NIKOLAI’s crosswalks are CASRAI’s own mappings of what these organizations have published, not endorsements by the organizations themselves, but they’re a starting point for comparing definitions side by side instead of taking any one framework’s usage as universal.

Frequently asked questions

Is there one official definition of “frontier AI model”?

No. At least three distinct approaches are in active use — compute thresholds (SB 53, the RAISE Act, the EU AI Act), capability thresholds (Anthropic’s Responsible Scaling Policy), and relative state-of-the-art definitions (the Frontier Model Forum) — and none of them takes precedence over the others. Which one applies depends on which law or framework you’re asking about.

Does a compute threshold mean a model is exempt forever if it stays under it?

Under SB 53 and the RAISE Act, no ongoing exemption works the other way either: once a developer’s model has met the 1026-operation threshold, that developer stays a “frontier developer” for that model even if later compute accounting would put it under the line. The threshold is a one-time gate, not a running eligibility test.

Can two frameworks disagree about whether the same model is “frontier”?

Yes, and this is expected rather than an error. A compute-threshold law and a capability-threshold framework are measuring different things, so a model can clear one line without clearing the other. See “Why the same model can be ‘frontier’ under one framework and not another” above.

Why does the EU AI Act have two different FLOPs numbers?

They answer two different questions. The 1023 FLOPs figure is a practical indicator for whether a model counts as “general-purpose” under the Act at all (Article 3(63)); the 1025 FLOPs figure is the higher bar that triggers a presumption of “systemic risk” under Article 51(1)(a), which is the tier that brings the heavier obligations.

Where can I see how these definitions map onto each other in detail?

Start with NIKOLAI, CASRAI’s dictionary of frontier-AI safety elements, and the individual framework guides in this cluster, including the SB 53 guide and the Responsible Scaling Policy guide.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about What Is a Frontier AI Model?

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →