Examples
Worked examples
- Is an instance
A model technical report including BIG-bench Hard average accuracy across 23 tasks.
- Is an instance
A new benchmark paper using BIG-bench as a baseline distribution of LLM capability.
Counter-examples
Looks similar, but isn't
- Not an instance
MMLU (a different benchmark).
- Not an instance
A single dataset like SQuAD.
Editorial commentary
BIG-bench (the Beyond the Imitation Game benchmark) is a community-contributed collection of more than 200 tasks, published in 2022 (Srivastava et al., arXiv:2206.04615), designed to probe language-model capabilities — reasoning, world knowledge, social bias, and more — that narrower, single-metric benchmarks tend to miss. Tasks were submitted and peer-reviewed by contributors across many institutions rather than authored by one lab.
Currency note: this benchmark is now largely historical
BIG-bench saturated as models improved: state-of-the-art systems now score near-ceiling on most of its original tasks, which is why a harder subset, BIG-Bench Hard (BBH, Suzgun et al., 2022), was carved out specifically from the tasks where models still underperformed humans. As of 2026, BBH itself has substantially saturated too — current frontier models score well above 0.9 on the public BBH leaderboard — prompting a further successor, BIG-Bench Extra Hard (BBEH, 2025), built by replacing each BBH task with a harder variant probing the same underlying reasoning capability; best-performing models on BBEH score roughly in the 10-45% range as of its introduction, restoring the discriminative power the original benchmark has lost. A page, procurement document, or grant proposal citing a raw BIG-bench (not BBH or BBEH) score as evidence of current model capability should be read with this saturation in mind.
How this differs from its siblings
See MLCommons benchmark for the fixed-workload, hardware-comparison alternative, and synthetic benchmark for benchmarks whose items are model-generated rather than human-authored (BIG-bench’s tasks were human-authored and peer-reviewed).
References
- Srivastava, A. et al. (2022). “Beyond the Imitation Game.” arXiv:2206.04615.
- Suzgun, M. et al. (2022). “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.”
- “BIG-Bench Extra Hard” (2025), arXiv:2502.19187.
- See also: MMLU benchmark, HELM benchmark.
Also known as
Beyond the Imitation Game Benchmark · BBH (subset)
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="BIG-bench"
vocab-term-identifier="https://casrai.org/dictionary/term/big-bench" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/big-bench",
"name": "BIG-bench",
"identifier": "https://casrai.org/dictionary/term/big-bench",
"description": "The Beyond the Imitation Game benchmark, a community-contributed collection of more than 200 tasks designed to probe capabilities of large language models that may be missed by narrower benchmarks.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
"url": "https://casrai.org/dictionary/term/big-bench",
"sameAs": [
"Beyond the Imitation Game Benchmark",
"BBH (subset)"
],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-05-21T02:22:51",
"dateModified": "2026-08-22T15:44:03",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}







