Direct comparison
Accession Number vs. DOI: Which to Cite
Accession numbers locate data in its home repository; DOIs are the resolvable ID for citation credit. When a dataset needs both, and which to cite where.
Written and maintained by CASRAI Editorial Board
Last updated
Ask CASRAI · included with Regulatory Radar
Ask about Accession Number vs. DOI: Which to Cite
Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.
150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.
Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.
How do Accession Number, DOI compare side by side?
The table below compares Accession Number, DOI across 15 procurement-relevant dimensions, from what it identifies through which to lead with in a citation.
Side-by-side comparison
| Dimension | Accession Number | DOI |
|---|---|---|
| What it identifies | A locally-scoped identifier assigned by a single domain-specific repository (e.g., NCBI GenBank, SRA, or GEO) to a specific record — a sequence, a sample, or a whole study/series — within that repository's own numbering system. | A globally unique, cross-domain identifier that can be assigned to almost any digital research output — articles, datasets, software, preprints — regardless of which repository or publisher hosts it. |
| Who assigns it | The domain repository itself, automatically, at the point of deposit — e.g., NCBI mints a GSE accession the moment a GEO series is submitted. No external registration agency is involved. | A DOI Registration Agency (RA) — for research data, almost always DataCite; for journal articles, almost always Crossref — under the DOI Foundation's shared technical standard (ISO 26324) and Handle System resolution infrastructure. The repository must be a DataCite member, or work through one, to mint DOIs. |
| Format | Repository-specific prefixes and patterns: GenBank sequences (e.g., AB123456.1, with a version suffix), SRA runs/experiments/studies (SRR/SRX/SRP/SRS + digits), GEO platforms/samples/series (GPL/GSM/GSE + digits). | A single universal syntax: a prefix assigned to the registrant (e.g., 10.5061) plus a suffix chosen by the registrant, e.g., 10.5061/dryad.abc123. The prefix format is identical across every domain and every registration agency. |
| Where it resolves | Only inside that repository's own search/retrieval system — searching a GEO accession on NCBI's site returns the record; the same string typically means nothing to a generic resolver or a different database. | Anywhere, via doi.org (built on the Handle System) — appending any DOI to https://doi.org/ redirects to the current location of the resource, independent of which organization currently hosts it. |
| Primary purpose | Retrieval and internal bookkeeping — locating, versioning, and cross-referencing the exact record inside the domain repository's own database and APIs (e.g., pulling raw reads by SRA accession). | Formal citation and credit — a stable reference a reader, indexer, or citation-tracking system can resolve and count, the same way it counts a journal-article citation. |
| Persistence commitment | Backed by the repository's own operational continuity — stable for decades in practice, but the guarantee is institutional, not contractual, and stops working if the repository is discontinued without a deliberate migration (see the ArrayExpress→BioStudies migration, where accessions were carried forward on purpose). | Backed by a contractual persistence obligation the registration agency and depositor accept as DataCite members — if the underlying resource moves, the DOI record's landing-page URL must be updated so the identifier keeps resolving, independent of any one repository's survival. |
| Required metadata to mint one | Whatever the repository's own submission requirements demand for that data type (e.g., GEO enforces MIAME/MINSEQE content categories: raw data, processed data, sample annotation, experimental design, platform annotation, protocols) — no universal minimum. | DataCite’s Mandatory Properties, the same six fields for every dataset DOI regardless of domain: Identifier, Creator, Title, Publisher, PublicationYear, ResourceType. A DataCite DOI cannot be minted without all six. |
| Versioning | Handled per repository convention — GenBank appends a version number to the accession itself (e.g., .1, .2); a GEO series is generally treated as a single evolving record rather than versioned releases. | Handled per registrant policy — many data repositories mint a new, version-specific DOI for each release plus a separate "concept DOI" that always resolves to the latest version, a pattern DataCite explicitly supports. |
| In a methods section | Cite the accession number to tell a reader exactly which record to pull for reproduction — it's the identifier someone actually types into the repository's search box or API. | Cite the DOI when the journal or repository has assigned one and the goal is a formally trackable, indexer-recognized citation — increasingly expected even for data a reader will still retrieve by accession. |
| In a data availability statement | Effectively required whenever the data lives in a domain repository — a statement that omits the accession number leaves a reader with no way to actually locate the deposited data. | Include it if the repository or a generalist archive assigned one — it gives the statement a permanently resolvable link and lets the dataset be formally counted as a citable output, but it does not replace the accession number for retrieval. |
| When a dataset has both | Common when primary data sits in a domain repository (a GEO series, an SRA run) while a companion DOI is minted separately — by a generalist repository holding a linked package, by a data-descriptor journal publishing alongside the deposit, or by a domain repository that itself participates in DataCite (as EMBL-EBI's BioStudies/ArrayExpress and the wwPDB do). | The DOI record's metadata typically lists the accession number as a relatedIdentifier, and the accession's own repository page typically links back out to the DOI — the two are meant to be used together, not as substitutes. |
| If the repository is discontinued | At risk — an accession is only as durable as the specific repository. When ArrayExpress migrated into BioStudies, its accession numbers were deliberately preserved as part of the migration precisely because this risk is real. | More resilient to any one repository's fate — because the landing page can be updated to point at wherever the data moved, the DOI keeps resolving even after a repository shuts down, provided the data was migrated somewhere DataCite-registered. |
| Cost to the depositor | Free — issuing accession numbers is a core, non-monetized function of domain repositories like GenBank, SRA, and GEO. | Usually absorbed by the repository rather than charged per deposit — DataCite charges member repositories an annual fee plus a small per-DOI fee; a repository must be a paying DataCite member (directly or via a consortium) to mint DOIs at all. |
| Typical repositories | GenBank, SRA, GEO, dbGaP, ProteomeXchange/PRIDE (PXD accessions), and most other domain-specific NIH/NCBI and EMBL-EBI databases. | Zenodo, Dryad, Figshare, OSF, Dataverse, and other DataCite-member generalist repositories by default; also assigned directly by some domain repositories, including EMBL-EBI's BioStudies and the wwPDB. |
| Which to lead with in a citation | Lead with the accession number when the citation's job is enabling retrieval within that domain's own tools and databases. | Lead with the DOI when the citation's job is formal scholarly credit and machine-trackable citation counting — cite both where both exist. |
Common questions
Common questions about Accession Number vs DOI
Can a dataset have both an accession number and a DOI?
+
Yes, and increasingly this is the norm rather than the exception. A dataset deposited in a domain repository like NCBI's GEO or SRA gets an accession number automatically; if that same dataset is also registered with DataCite — directly by the repository, via a linked generalist-repository deposit, or through a data-descriptor publication — it also gets a DOI. The two identifiers point at the same underlying data and are meant to be cited together, not as alternatives.
Which one should I cite in my methods section?
+
Cite the accession number if a reader's next step is to actually retrieve the data from its domain repository — that's the identifier the repository's own search box and API expect. If the deposit also has a DOI, cite both: the accession for retrieval, the DOI for a formally trackable citation.
Which one should I cite in a data availability statement?
+
Include the accession number whenever the data lives in a domain repository — without it, the statement doesn't actually tell a reader how to find the data. Add the DOI alongside it if one was assigned; it strengthens the statement without replacing the retrieval-specific accession number.
Why are DOIs increasingly preferred for citation credit if accession numbers already work?
+
Because DOIs are built for citation tracking across the whole scholarly ecosystem, not just retrieval within one repository. A DOI resolves through the universal doi.org/Handle System regardless of which organization currently hosts the data, carries mandatory DataCite metadata (creator, title, publisher, year, resource type) that citation managers and indexers can parse automatically, and lets a dataset accumulate formally countable citations the way a journal article does. An accession number was never designed to do any of that — it was designed to let a repository's own tools find a specific record.
Does an accession number ever resolve outside its own repository?
+
Generally no. An accession number is meaningful within the numbering scheme of the repository that issued it — a GEO GSE number typed into a generic search engine or a different database usually returns nothing useful unless that system has specifically built in GEO-accession recognition. This is the core practical difference from a DOI, which is designed to resolve identically everywhere.
What happens to an accession number if the repository is discontinued or migrates?
+
It depends entirely on the migration. When EMBL-EBI folded ArrayExpress into BioStudies, the original accession numbers were deliberately carried forward so existing citations kept working — but that continuity is a deliberate choice by the repository, not a guarantee built into the accession system itself. A DOI's persistence, by contrast, is a contractual obligation the registration agency and depositor accept.
Is a DOI more 'official' or trustworthy than an accession number?
+
No — they answer different questions and neither outranks the other. An accession number from a recognized domain repository like GenBank or SRA is exactly as authoritative for locating the deposited data as it has ever been; a DOI adds resolvability and formal citation infrastructure on top, it does not correct or supersede the accession.
Going deeper








