Examples
Worked examples
- Is an instance
BASE (Bielefeld Academic Search Engine) aggregating from thousands of repositories worldwide.
- Is an instance
OpenAlex aggregating publication metadata from Crossref, ORCID, ROR, and crawled sources.
Counter-examples
Looks similar, but isn't
- Not an instance
A single repository serving only its own holdings is not an aggregator.
- Not an instance
A search engine that indexes general web content (without harvesting structured metadata) is not, strictly, a research-information aggregator.
Editorial commentary
An aggregator service sits downstream of repositories and CRIS systems: it harvests metadata (and sometimes full text) from many independent sources, performs entity deduplication and enrichment, and exposes the combined corpus through one search interface or API. Aggregators are what make thousands of independently-run repositories visible and searchable as one corpus.
Named examples and how they differ
- OpenAIRE — EU-funded, strong emphasis on funder open-access compliance monitoring alongside discovery.
- BASE (Bielefeld Academic Search Engine) — a long-running academic-search aggregator operated by Bielefeld University Library.
- CORE — operated by the Open University/Jisc, with a focus on aggregating open-access full text for text-and-data-mining.
- OpenAlex — a large-scale scholarly metadata graph, itself a POSI signatory, reaffirmed after the 2025 POSI v2.0 update.
How aggregators harvest data
Aggregators pull source data via OAI-PMH, ResourceSync, REST APIs, and bulk data dumps. This is the same protocol machinery an open archive exposes to be harvestable — the two concepts are complementary: an open archive is the compliant source, an aggregator is the downstream service that combines results across many such sources.
Dependency on upstream data quality
Aggregation is only as good as the metadata it harvests. Entity resolution — identifying that two records from different sources describe the same author, organisation or output — is a genuinely hard, ongoing problem; persistent identifiers such as ORCID iDs, ROR IDs and DOIs materially reduce it by giving aggregators an unambiguous match key instead of fuzzy name matching.
How aggregators differ from repositories, CRIS, and registries
A repository is a primary storage location where content originates. A CRIS is an institution’s own system of record for its research activity. A registry (ROR, re3data) is a curated list of entities. An aggregator holds none of these as its origin — it is a derived index built by harvesting from many such sources, and generally does not accept direct researcher deposits.
Frequently asked questions
Is OpenAlex an aggregator or a database?
Functionally both — it aggregates metadata from many upstream sources and exposes the combined result as a queryable database and API.
Do aggregators host full text?
Some do (CORE emphasises this), but many aggregate metadata only and link back to the source repository.
How current is aggregator data versus the source?
Depends on harvest frequency — most re-harvest daily to weekly, so there is typically some lag from the source update.
References
- Knoth P, Pontika N, ‘Aggregating Research Papers from Publishers’ Systems to Support Text and Data Mining’, CORE/Open University working papers. OpenAlex POSI signatory status (openscholarlyinfrastructure.org, accessed 2026-08-22).
Also known as
Metadata aggregator · Discovery service
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="Aggregator service"
vocab-term-identifier="https://casrai.org/dictionary/term/aggregator-service" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/aggregator-service",
"name": "Aggregator service",
"identifier": "https://casrai.org/dictionary/term/aggregator-service",
"description": "A service that harvests, harmonises, and re-exposes metadata and (sometimes) content from many upstream sources, providing a unified search, browse, or query interface across the aggregated corpus; canonical examples include OpenAIRE, BASE, CORE, and OpenAlex.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
"url": "https://casrai.org/dictionary/term/aggregator-service",
"sameAs": [
"Metadata aggregator",
"Discovery service"
],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-05-21T02:22:48",
"dateModified": "2026-08-22T15:38:25",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}







