Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

Datasheet for datasets

A structured document accompanying a machine-learning dataset that records its motivation, composition, collection process, pre-processing, intended uses, distribution, and maintenance, modelled on electronic-component datasheets.

ByCASRAI Editorial Board
· Last updated 15 Aug 2026
Share this

Ask CASRAI · included with Regulatory Radar

Ask about Datasheet for datasets

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A face-recognition benchmark distributed with a datasheet listing demographic composition, consent procedures, and recommended uses.

  • Is an instance

    An NLP corpus accompanied by a datasheet documenting source domains and crawling rules.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A dataset README containing only file format and column descriptions.

  • Not an instance

    A model card (covers the model, not the dataset).

Editorial commentary

Gebru et al. (2021) proposed datasheets to make dataset provenance and limitations visible to downstream model builders. Topics covered include consent and licensing of subjects, sampling and labelling procedures, demographic composition, known biases, and recommended/cautioned uses. Datasheets are complementary to model cards.

References

  • Gebru et al., 'Datasheets for datasets' (Communications of the ACM, 2021).

Frequently Asked Questions

Who introduced datasheets for datasets, and where was the concept published?

The concept was proposed by Gebru et al. in a 2021 paper, ‘Datasheets for datasets,’ published in Communications of the ACM, to make dataset provenance and limitations visible to downstream model builders.

What does a datasheet for datasets document?

It is a structured document accompanying a machine-learning dataset that records its motivation, composition, collection process, pre-processing, intended uses, distribution, and maintenance, modelled on datasheets used for electronic components. In practice this covers consent and licensing of subjects, sampling and labelling procedures, demographic composition, known biases, and recommended or cautioned uses.

How is a datasheet for datasets different from a model card?

A datasheet documents the dataset itself, while a model card documents the model trained on it. The two are complementary, not interchangeable.

Does a dataset README count as a datasheet for datasets?

No. A README that only lists file formats and column descriptions does not qualify. A datasheet for datasets additionally covers consent, sampling and labelling procedures, demographic composition, known biases, and recommended or cautioned uses.

What’s an example of a dataset with a datasheet for datasets?

Examples include a face-recognition benchmark distributed with a datasheet listing its demographic composition, consent procedures, and recommended uses, and an NLP corpus accompanied by a datasheet documenting its source domains and crawling rules.

Also known as

dataset datasheet

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Datasheet for datasets"
      vocab-term-identifier="https://casrai.org/dictionary/term/datasheet-for-datasets" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/datasheet-for-datasets",
  "name": "Datasheet for datasets",
  "identifier": "https://casrai.org/dictionary/term/datasheet-for-datasets",
  "description": "A structured document accompanying a machine-learning dataset that records its motivation, composition, collection process, pre-processing, intended uses, distribution, and maintenance, modelled on electronic-component datasheets.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/datasheet-for-datasets",
  "sameAs": [
    "dataset datasheet"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:50",
  "dateModified": "2026-08-15T04:41:15",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.