Skip to main content
v2026.11,858 entries · CC-BY 4.0

Direct comparison

What Gets Graded: Model, Company, or Document

AILuminate grades a deployed system, the FLI Index grades a company, SaferAI grades a framework document. Why their rankings disagree.

Written and maintained by CASRAI Editorial Board

Last updated

Ask CASRAI · free to try

Ask about What Gets Graded: Model, Company, or Document

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

How do MLCommons AILuminate, FLI AI Safety Index, SaferAI Frontier Risk Management Tracker compare side by side?

The table below compares MLCommons AILuminate, FLI AI Safety Index, SaferAI Frontier Risk Management Tracker across 10 procurement-relevant dimensions, from what carries the grade through what it is honestly good for.

Side-by-side comparison

DimensionMLCommons AILuminateFLI AI Safety IndexSaferAI Frontier Risk Management Tracker
What carries the gradeA system-under-test (SUT): one fixed, complete chatbot configuration from prompt in to response out. It may wrap several models, guardrail layers and retrieval steps, or none. The benchmark separately lists productised AI systems and bare models, and the same underlying model can appear in both lists with different grades.A company. The Summer 2026 edition grades nine firms: Anthropic, OpenAI, Google DeepMind, Meta, Z.ai, Alibaba Cloud, xAI, DeepSeek and Mistral. No individual model receives a grade.A single published document: the company's frontier safety framework. The tracker's methodology page is explicit that it “exclusively assess[es] companies' frontier safety frameworks”, a narrowing from its first iteration, which read all publicly available documents.
Who does the gradingMLCommons' AI Risk and Reliability working group, using an ensemble evaluator model built from open safety evaluators and fine-tuned for grading confidence. The scoring is automated; MLCommons says it analyses results to keep testing as fair as possible.Seven independent reviewers, named in the report: David Krueger, Sharon Li, Tegan Maharaj, Sneha Revanur, Stuart Russell, Robert Trager and Yi Zeng. Human expert judgement against absolute standards, with discretionary weighting.SaferAI's own analysts. Every criterion score carries a written rationale quoting the company's framework with page references, and scores are peer-reviewed among the authors before publication.
Evidence the grader readsModel outputs. Over 24,000 test prompts per language, split into a 12,000-prompt public practice set and a 12,000-prompt private official set, across twelve hazard categories grouped as physical, non-physical and contextual hazards. A jailbreak suite exists separately at v0.5 draft status.Published policies, research, reporting and company disclosures, supplemented by a survey sent to each company. Companies that do not return the survey are graded on the public record alone.The text of the framework document, nothing else. Not model behaviour, not internal practice, not what the company says elsewhere.
Scale, and what the top score meansFive bands: Poor, Fair, Good, Very Good, Excellent. Four of them are relative to a reference system (Poor is more than 3x its violation rate; Very Good is under 0.5x). Only Excellent is absolute: under 0.1 percent violating responses.US GPA letters, A (4.0) through F (0). Grades are assigned against absolute standards, not a curve, which is why the whole field can score badly at once. In Summer 2026 the best overall grade was a C+.A 0-100 percent score per criterion on deliberately non-uniform steps: 0, 10, 25, 50, 75, 90, 100, finer at the extremes and coarser in the middle. Four dimensions weighted equally at 25 percent each.
The published resultsPublic grades for every English and French system-under-test, both overall and per hazard category. Ten AI systems and seventeen bare models appear in the v1.0 English results.Anthropic C+ (2.66), OpenAI C (2.28), Google DeepMind C (2.01), Meta D+ (1.32), Z.ai D- (0.88), Alibaba Cloud D- (0.87), xAI F (0.65), DeepSeek F (0.47), Mistral F (0.33).Anthropic 35, OpenAI 34, Microsoft 33, Meta 33, G42 24, Google DeepMind 20, xAI 18, Amazon 18, NVIDIA 16, Magic 11, Naver 10, Cohere 8. The twelve scores average about 22 percent.
Who is in scope, and how you get inVendor-facing. For the v1.0 launch MLCommons says it “selected the vendors of the greatest public interest globally or regionally” and worked with them to pick one cutting-edge and one accessible model each. Sponsors were offered the option to appear in public results.Reviewer-selected. The nine companies were chosen as the frontier developers of interest; participation is not required for a company to be graded.Any company that has published a frontier safety framework. That is why Microsoft, NVIDIA, Amazon, Cohere, Naver, G42 and Magic appear here but not in the FLI table, and why Z.ai and Alibaba Cloud appear in the FLI table but not here.
Can the subject declineEffectively yes. The public v1.0 results mark three systems as opted out: Grok-3-Preview, Hunyuan-TurboS and Llama 3.3 49b Nemotron Super. An absent grade is therefore not a bad grade, and should never be read as one.Not from the grading, only from the survey. A company that ignores the survey is still graded, on whatever public record exists, which structurally penalises firms with strong internal but unpublished practice.Only by not publishing a framework at all. Once a framework is public, it is gradeable without the company’s cooperation, since the input is a document already in the open.
How current a score staysVersion-bound. Every official score is tied to a specific release version of the benchmark, so a grade from one release is not directly comparable to a grade from another. The arXiv paper describing v1.0 was posted in February 2025 and revised that April; a v1.1 suite now exists.Edition-bound, published roughly twice a year. Winter 2025, Summer 2025 and Summer 2026 editions all exist, so a quoted grade needs its edition named or it means very little.Snapshot-bound. The current comparison table is marked up to date as of July 2026. The methodology is deliberately held fixed over time so scores stay comparable across snapshots.
What the grade explicitly does not coverGovernance. The FAQ is direct that governance practices are not evaluated and that a good grade means the system “presents as low or lower risk”, “not that it is risk free”. The benchmark claims only negative predictive power: “Performing well on the benchmark does not mean that your model is safe, simply that we have not identified critical safety weaknesses.” Known limits include single-turn-only interaction, limited language coverage and evaluator uncertainty.Model behaviour. Nothing in the Index is a measurement of what a deployed system does when prompted. The report also flags that comparing Chinese and US firms fairly is difficult because their regulatory contexts differ.Whether any of it is true in practice. A framework can score well because it is written thoroughly and still describe procedures the company does not follow; the tracker reads text, not conduct.
What it is honestly good forScreening a specific configured deployment for hazard-category behaviour before you put it in front of users, and comparing candidate configurations against each other on the same benchmark release.A structured external read on a vendor as an organisation: does it publish, does it evaluate, does it share information, does it have accountable governance. Useful as one input to vendor due diligence.Checking whether a vendor’s framework actually contains the risk-identification, analysis, treatment and governance commitments you are about to rely on contractually, with page-level citations you can follow.

Common questions

Common questions about MLCommons AILuminate vs FLI AI Safety Index vs SaferAI Frontier Risk Management Tracker

Can I use these three to rank the same set of vendors?

+

No, and the overlap is smaller than it looks. FLI grades nine companies and SaferAI grades twelve, but only five sit on both lists: Anthropic, OpenAI, Google DeepMind, Meta and xAI. DeepSeek, Mistral, Z.ai and Alibaba Cloud are graded by FLI and absent from SaferAI, which instead covers Microsoft, Amazon, NVIDIA, Cohere, Naver, G42 and Magic. AILuminate ranks neither list, because it ranks configured chatbots. If you need one table covering one set of vendors, you have to build it yourself and label which grader supplied which column.

A vendor scores well on one and badly on another. Which one is right?

+

Both, usually. Meta at 33 percent on SaferAI and D+ on the FLI Index is the clearest case in the current data. SaferAI read Meta’s published frontier safety framework and found it drafted roughly as thoroughly as Anthropic’s. FLI read Meta’s whole public record across risk assessment, current harms, safety frameworks, existential safety, governance and information sharing, and found most of those domains weak. A well-written framework document inside a company with a poor overall disclosure record produces exactly this pattern, and noticing it is more informative than either number alone.

Does a good AILuminate grade mean the model is safe?

+

No, and MLCommons says so in its own materials: “Performing well on the benchmark does not mean that your model is safe, simply that we have not identified critical safety weaknesses.” The benchmark is designed to have negative predictive power only. Four of its five grade bands are defined relative to a reference system rather than to an absolute violation rate, results carry considerable variance from prompt selection and automated evaluation, and the current suite is single-turn and English-and-French-first. A grade tells you a weakness was not found in that configuration under those prompts.

Which one should a university procurement or research computing office look at?

+

It depends on the question. If you are standing up a campus AI assistant, the object you care about is your own deployment, which is its own system-under-test: the vendor’s bare-model grade does not transfer to your configuration once you add guardrails, a retrieval layer and a system prompt, and AILuminate’s own SUT definition is the reason why. If the question is instead whether the vendor is an organisation you can rely on across a multi-year agreement, that is an FLI and SaferAI question. A sponsored programs or research security review asking about governance, disclosure and incident handling gets nothing useful from a hazard benchmark, and an IRB asking whether a study tool will hand participants unqualified medical or legal advice gets nothing useful from a company letter grade.

Do any of these grades have legal force?

+

None of the three. All are voluntary third-party assessments by non-governmental organisations, and no regulator adopts their scores. What regulation does require, in places, is the underlying artefact rather than the grade: California SB 53 and the EU AI Act’s general-purpose AI obligations put weight on a developer publishing and maintaining its own framework and reporting. That makes SaferAI’s object, the framework document, the one that most closely tracks what a statute already expects to exist, without making SaferAI’s percentage a compliance finding.

Why does SaferAI only read the framework document?

+

Because it makes the comparison auditable. The tracker applies 65 criteria across four equally weighted dimensions, and attaches to each score a rationale quoting the company’s framework with page numbers, so an outside reader can check the scoring rather than trust it. That is only possible against a fixed public text. The cost of that choice is the obvious one, and SaferAI does not hide it: the tracker measures what a company has written down, not what it does.

Do any of them check whether the evaluation itself was valid?

+

Not directly, and it is the largest shared gap. A model can behave differently when it detects it is being evaluated, can underperform deliberately, or can be optimised against a known grader. CASRAI’s NIKOLAI dictionary carries a controlled list for exactly these conditions under its N5 evidence and evaluations track, the evaluation-validity threat element, naming evaluation awareness, sandbagging, alignment faking, metagaming or grader-gaming, and reward hacking. NIKOLAI is CASRAI’s own independent dictionary and is not endorsed by MLCommons, the Future of Life Institute or SaferAI; any alignment between its terms and theirs is a shadow mapping unless one of those organisations files a Mapping Declaration.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

Research-admin question? Get an answer that links its sources.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.