Examples
Worked examples
- Is an instance
NCI Genomic Data Commons (gdc.cancer.gov) hosting TCGA and many other cancer-genomic datasets.
- Is an instance
AnVIL (NHGRI Analysis Visualization and Informatics Lab-space) for human-genomic data analysis.
Counter-examples
Looks similar, but isn't
- Not an instance
A read-only data dump on a website is not a data commons.
- Not an instance
A single-institution dataset with no shared tooling or governance is not a data commons.
Editorial commentary
A data commons is a research-data infrastructure that co-locates a curated dataset (or federation of datasets) with the compute and tooling needed to analyse it, under an explicit governance body and a contribution/access policy — rather than simply hosting files for download. Grossman et al.’s 2016 ‘A Case for Data Commons’ names four defining components: a curated data set (or federation of sets); shared compute and tooling co-located with the data so analysis happens near the data rather than after a bulk download; a governance body responsible for the commons’ operation; and a contribution/access policy balancing inclusion and reuse against protecting sensitive data and contributor rights. The term draws on Elinor Ostrom’s work on commons governance, adapted to a research-data context, and is widely used in biomedical and earth-science research infrastructure.
Examples
The NIH National Cancer Institute’s Genomic Data Commons (GDC), the NIH Common Fund Data Ecosystem (CFDE), and AnVIL (a cloud-based genomic-analysis commons funded by NHGRI) are commonly cited exemplars — each pairs a specific curated dataset with dedicated compute and an explicit governance and access-policy layer.
How this differs from a data hub, a data lake, and national data infrastructure
A data hub aggregates and harmonises data from multiple upstream sources into one opinionated model, but does not require co-located compute or a formal governance body as a defining feature. A data lake is a raw, schema-on-read storage pattern with no governance-body or access-policy requirement built into the definition at all. National data infrastructure operates one level up again: it names a funder- or government-level coordinating programme (such as the UK Data Service or Germany’s NFDI) that may fund or operate several data commons, hubs, or repositories, rather than being an architecture pattern for a single dataset ecosystem itself. A data commons is the specific combination of curated data, co-located compute, governance, and access policy around one dataset or federation; the other three terms name different, adjacent patterns.
References
- Grossman R.L. et al., ‘A Case for Data Commons: Toward Data Science as a Service’, Computing in Science & Engineering 18(5), 2016.
Also known as
Research data commons
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="Data commons"
vocab-term-identifier="https://casrai.org/dictionary/term/data-commons" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/data-commons",
"name": "Data commons",
"identifier": "https://casrai.org/dictionary/term/data-commons",
"description": "A shared data resource — often combined with shared computing and analysis tools — governed by a community under defined access and contribution rules, designed to enable many users to use and add to the resource for collective benefit.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/data-commons",
"sameAs": [
"Research data commons"
],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-05-21T02:17:49",
"dateModified": "2026-08-23T04:44:18",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}







