Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack BStablev2026.2

Aggregator service

A service that harvests, harmonises, and re-exposes metadata and (sometimes) content from many upstream sources, providing a unified search, browse, or query interface across the aggregated corpus; canonical examples include OpenAIRE, BASE, CORE, and OpenAlex.

ByCASRAI Editorial Board
· Last updated 22 Aug 2026
Share this

Ask CASRAI · included with Regulatory Radar

Ask about Aggregator service

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    BASE (Bielefeld Academic Search Engine) aggregating from thousands of repositories worldwide.

  • Is an instance

    OpenAlex aggregating publication metadata from Crossref, ORCID, ROR, and crawled sources.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A single repository serving only its own holdings is not an aggregator.

  • Not an instance

    A search engine that indexes general web content (without harvesting structured metadata) is not, strictly, a research-information aggregator.

Editorial commentary

An aggregator service sits downstream of repositories and CRIS systems: it harvests metadata (and sometimes full text) from many independent sources, performs entity deduplication and enrichment, and exposes the combined corpus through one search interface or API. Aggregators are what make thousands of independently-run repositories visible and searchable as one corpus.

Named examples and how they differ

  • OpenAIRE — EU-funded, strong emphasis on funder open-access compliance monitoring alongside discovery.
  • BASE (Bielefeld Academic Search Engine) — a long-running academic-search aggregator operated by Bielefeld University Library.
  • CORE — operated by the Open University/Jisc, with a focus on aggregating open-access full text for text-and-data-mining.
  • OpenAlex — a large-scale scholarly metadata graph, itself a POSI signatory, reaffirmed after the 2025 POSI v2.0 update.

How aggregators harvest data

Aggregators pull source data via OAI-PMH, ResourceSync, REST APIs, and bulk data dumps. This is the same protocol machinery an open archive exposes to be harvestable — the two concepts are complementary: an open archive is the compliant source, an aggregator is the downstream service that combines results across many such sources.

Dependency on upstream data quality

Aggregation is only as good as the metadata it harvests. Entity resolution — identifying that two records from different sources describe the same author, organisation or output — is a genuinely hard, ongoing problem; persistent identifiers such as ORCID iDs, ROR IDs and DOIs materially reduce it by giving aggregators an unambiguous match key instead of fuzzy name matching.

How aggregators differ from repositories, CRIS, and registries

A repository is a primary storage location where content originates. A CRIS is an institution’s own system of record for its research activity. A registry (ROR, re3data) is a curated list of entities. An aggregator holds none of these as its origin — it is a derived index built by harvesting from many such sources, and generally does not accept direct researcher deposits.

Frequently asked questions

Is OpenAlex an aggregator or a database?

Functionally both — it aggregates metadata from many upstream sources and exposes the combined result as a queryable database and API.

Do aggregators host full text?

Some do (CORE emphasises this), but many aggregate metadata only and link back to the source repository.

How current is aggregator data versus the source?

Depends on harvest frequency — most re-harvest daily to weekly, so there is typically some lag from the source update.

References

  • Knoth P, Pontika N, ‘Aggregating Research Papers from Publishers’ Systems to Support Text and Data Mining’, CORE/Open University working papers. OpenAlex POSI signatory status (openscholarlyinfrastructure.org, accessed 2026-08-22).

Also known as

Metadata aggregator · Discovery service

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Aggregator service"
      vocab-term-identifier="https://casrai.org/dictionary/term/aggregator-service" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/aggregator-service",
  "name": "Aggregator service",
  "identifier": "https://casrai.org/dictionary/term/aggregator-service",
  "description": "A service that harvests, harmonises, and re-exposes metadata and (sometimes) content from many upstream sources, providing a unified search, browse, or query interface across the aggregated corpus; canonical examples include OpenAIRE, BASE, CORE, and OpenAlex.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/data-infrastructure#set",
  "url": "https://casrai.org/dictionary/term/aggregator-service",
  "sameAs": [
    "Metadata aggregator",
    "Discovery service"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:48",
  "dateModified": "2026-08-22T15:38:25",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.