Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

BIG-bench

The Beyond the Imitation Game benchmark, a community-contributed collection of more than 200 tasks designed to probe capabilities of large language models that may be missed by narrower benchmarks.

ByCASRAI Editorial Board
· Last updated 22 Aug 2026
Share this

Ask CASRAI · included with Regulatory Radar

Ask about BIG-bench

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A model technical report including BIG-bench Hard average accuracy across 23 tasks.

  • Is an instance

    A new benchmark paper using BIG-bench as a baseline distribution of LLM capability.

Counter-examples

Looks similar, but isn't

  • Not an instance

    MMLU (a different benchmark).

  • Not an instance

    A single dataset like SQuAD.

Editorial commentary

BIG-bench (the Beyond the Imitation Game benchmark) is a community-contributed collection of more than 200 tasks, published in 2022 (Srivastava et al., arXiv:2206.04615), designed to probe language-model capabilities — reasoning, world knowledge, social bias, and more — that narrower, single-metric benchmarks tend to miss. Tasks were submitted and peer-reviewed by contributors across many institutions rather than authored by one lab.

Currency note: this benchmark is now largely historical

BIG-bench saturated as models improved: state-of-the-art systems now score near-ceiling on most of its original tasks, which is why a harder subset, BIG-Bench Hard (BBH, Suzgun et al., 2022), was carved out specifically from the tasks where models still underperformed humans. As of 2026, BBH itself has substantially saturated too — current frontier models score well above 0.9 on the public BBH leaderboard — prompting a further successor, BIG-Bench Extra Hard (BBEH, 2025), built by replacing each BBH task with a harder variant probing the same underlying reasoning capability; best-performing models on BBEH score roughly in the 10-45% range as of its introduction, restoring the discriminative power the original benchmark has lost. A page, procurement document, or grant proposal citing a raw BIG-bench (not BBH or BBEH) score as evidence of current model capability should be read with this saturation in mind.

How this differs from its siblings

See MLCommons benchmark for the fixed-workload, hardware-comparison alternative, and synthetic benchmark for benchmarks whose items are model-generated rather than human-authored (BIG-bench’s tasks were human-authored and peer-reviewed).

References

  • Srivastava, A. et al. (2022). “Beyond the Imitation Game.” arXiv:2206.04615.
  • Suzgun, M. et al. (2022). “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.”
  • “BIG-Bench Extra Hard” (2025), arXiv:2502.19187.
  • See also: MMLU benchmark, HELM benchmark.

Also known as

Beyond the Imitation Game Benchmark · BBH (subset)

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="BIG-bench"
      vocab-term-identifier="https://casrai.org/dictionary/term/big-bench" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/big-bench",
  "name": "BIG-bench",
  "identifier": "https://casrai.org/dictionary/term/big-bench",
  "description": "The Beyond the Imitation Game benchmark, a community-contributed collection of more than 200 tasks designed to probe capabilities of large language models that may be missed by narrower benchmarks.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/big-bench",
  "sameAs": [
    "Beyond the Imitation Game Benchmark",
    "BBH (subset)"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-08-22T15:44:03",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.