Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

MLCommons benchmark

A benchmark published by the MLCommons consortium for measuring AI system performance under standardised workloads, datasets, and submission rules, with the principal suites being MLPerf Training, MLPerf Inference, and MLPerf HPC.

ByCASRAI Editorial Board
· Last updated 22 Aug 2026
Share this

Ask CASRAI · included with Regulatory Radar

Ask about MLCommons benchmark

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A vendor's MLPerf Inference v4.0 submission for a server-class GPU.

  • Is an instance

    A research group's open-division submission demonstrating a novel system architecture.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A single-paper benchmark report not subject to peer-submission review.

  • Not an instance

    A leaderboard maintained on a personal blog.

Editorial commentary

An MLCommons benchmark is a benchmark published by the MLCommons consortium — a nonprofit, multi-organisation industry-and-academic collaboration — for measuring AI system performance under standardised workloads, fixed datasets, and published submission rules that all participants must follow. The principal, longest-running suites are MLPerf Training (time/throughput to train a reference model to a target accuracy), MLPerf Inference (latency and throughput serving a trained model), and MLPerf HPC (training at supercomputer scale); MLCommons has more recently added AILuminate, a benchmark focused on AI-system safety rather than raw performance, assessing model responses against defined hazard categories.

How this differs from BIG-bench and a synthetic benchmark

MLCommons benchmarks are fixed-workload, standardised-hardware comparisons — their purpose is letting different vendors’ systems be compared on exactly the same task under exactly the same rules, closer in spirit to an industry conformance test than a capability probe. BIG-bench is a community-crowdsourced collection of capability-probing tasks with no fixed hardware-comparison framing. A synthetic benchmark is defined by how its test items were generated (by a model or procedure, rather than curated from real submissions), a dimension MLCommons benchmarks are largely orthogonal to, since MLPerf tasks use real reference datasets and models.

Why it matters for procurement

A research-computing procurement citing “MLPerf-benchmarked” hardware is citing a specific, rule-bound, third-party-audited comparison — a materially stronger claim than an unqualified vendor performance number.

References

Also known as

MLPerf

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="MLCommons benchmark"
      vocab-term-identifier="https://casrai.org/dictionary/term/mlcommons-benchmark" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/mlcommons-benchmark",
  "name": "MLCommons benchmark",
  "identifier": "https://casrai.org/dictionary/term/mlcommons-benchmark",
  "description": "A benchmark published by the MLCommons consortium for measuring AI system performance under standardised workloads, datasets, and submission rules, with the principal suites being MLPerf Training, MLPerf Inference, and MLPerf HPC.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/mlcommons-benchmark",
  "sameAs": [
    "MLPerf"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-08-22T15:44:02",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.