Written and maintained by CASRAI Editorial Board
Last updated
The h-index definition and its known limitations are covered on this site’s h-index dictionary entry. This guide is the practical companion to that entry: it walks through the calculation mechanics in more depth, works a second numeric example that includes a tie in citation counts, and focuses specifically on why the same researcher can get three different h-index numbers from Google Scholar, Scopus, and Web of Science — and what to do about that when reporting the figure.
Hirsch’s original definition
The h-index was proposed by physicist Jorge E. Hirsch in a 2005 paper, “An index to quantify an individual’s scientific research output,” published in Proceedings of the National Academy of Sciences (PNAS 102(46):16569–16572). Hirsch’s definition: a researcher has an h-index of h if h of their papers have each been cited at least h times, while the remaining papers have each been cited fewer than h times. The index was designed as a single number balancing productivity (how many papers) against impact (how often they’re cited), specifically to resist being distorted either by a large volume of rarely-cited papers or by one or two extremely highly cited outliers.
How to calculate an h-index manually
- Pick one citation database and stay inside it. Google Scholar, Scopus, and Web of Science each index a different set of documents and count citations differently (see below), so mixing citation counts pulled from different sources into one list will produce a number that doesn’t correspond to any real database’s h-index.
- List every paper with its total citation count from that database, for the researcher (or, for a journal or research group, the same method applies to that entity’s output).
- Sort the list by citation count, highest to lowest, and assign each paper a rank starting at 1.
- Compare each paper’s citation count to its rank, moving down the sorted list.
- Find the last rank position where the citation count is still greater than or equal to the rank number. That rank number is the h-index. As soon as you reach a rank where the citation count drops below the rank, the h-index has already been passed — it’s the rank immediately before that point.
Most researchers never do this by hand for a large publication list — Google Scholar Citations profiles compute and display it automatically, and Scopus/Web of Science author profiles do the same for their own indexed content (see the “finding it automatically” section below). The manual method matters for understanding what the displayed number actually represents, and for computing it yourself from a citation export when a profile isn’t available or needs checking.
Worked example
Take a researcher with ten papers, with these total citation counts already sorted highest to lowest: 55, 30, 30, 18, 12, 7, 7, 3, 1, 0.
| Rank | Citations | Citations ≥ rank? |
|---|---|---|
| 1 | 55 | Yes |
| 2 | 30 | Yes |
| 3 | 30 | Yes |
| 4 | 18 | Yes |
| 5 | 12 | Yes |
| 6 | 7 | Yes (7 ≥ 6) |
| 7 | 7 | Yes (7 ≥ 7) |
| 8 | 3 | No (3 < 8) |
| 9 | 1 | No |
| 10 | 0 | No |
Rank 7 is the last position where the condition still holds (7 citations at rank 7), and rank 8 fails it (3 citations is less than 8). This researcher’s h-index is 7. Note the tied citation count at ranks 2 and 3 (30 citations each) doesn’t need special handling — ties simply occupy adjacent ranks in whichever order, and the count-versus-rank comparison works the same way regardless of which tied paper is listed first. This example uses illustrative numbers only, constructed to show the mechanics, not a real researcher’s publication record.
Why Google Scholar, Scopus, and Web of Science give different h-index values for the same person
It’s normal, not an error, for the same researcher to have three different h-index numbers depending on which database computed it. The underlying cause in all three cases is coverage scope — which documents each database indexes, and which of those documents’ citations it counts:
- Google Scholar is an automated web crawl with no published inclusion criteria: it indexes journal articles, conference papers, theses, preprints, books and book chapters, patents, and other grey literature, and counts citations from any of that indexed material, including non-peer-reviewed sources. This breadth means Google Scholar h-index values are typically the highest of the three for the same person — comparative bibliometric studies have found Google Scholar h-index figures running roughly 1.3–1.4x higher than the equivalent Web of Science figure on average, though the gap varies substantially by field and career stage.
- Scopus (Elsevier) is curated against published selection criteria by its Content Selection and Advisory Board (CSAB) and counts citations only from Scopus-indexed source documents — a narrower, vetted set than Google Scholar’s crawl but broader in journal count than Web of Science’s Core Collection in most subject areas. Scopus launched in November 2004, and its citation counting is naturally shallower for citations that predate its own indexed backfile in a given subject area, which can undercount total citations (and therefore h-index) for researchers with much older highly-cited work. Scopus h-index values tend to fall between Web of Science and Google Scholar, typically modestly above Web of Science for the same person.
- Web of Science (Clarivate) draws on the curated Core Collection — Science Citation Index Expanded, Social Sciences Citation Index, Arts & Humanities Citation Index, Emerging Sources Citation Index, plus the Conference Proceedings and Book Citation Indexes — each with its own editorial selection criteria and, in several of those indexes, a much longer backfile (the Science Citation Index’s coverage reaches back to 1945, Social Sciences Citation Index to 1956). Its citation counts only include citations made by other Web of Science–indexed documents, which is generally the narrowest of the three counting scopes for current researchers, and typically produces the lowest h-index of the three for the same person.
None of the three numbers is “wrong” — they’re each an accurate count within that database’s own coverage. This is exactly why responsible-assessment frameworks such as DORA and CoARA caution against citing an h-index without naming the source database and pull date, and against comparing h-index figures pulled from different databases as if they were the same measurement. See the h-index inflation entry for a related distortion — how large-team, hyperauthorship papers can inflate an h-index within any of the three databases.
Finding your h-index automatically instead of calculating it by hand
- Google Scholar Citations profile computes and displays h-index (and i10-index) automatically once a profile is set up and publications are claimed — see this site’s guide to creating and optimizing a Google Scholar profile.
- Scopus Author Details pages display an author-level h-index computed from Scopus-indexed documents, reachable once an author has a disambiguated Scopus Author ID — see how to create a Scopus Author ID.
- Web of Science author/Researcher Profile pages display the same figure computed from Core Collection content — see setting up a Web of Science researcher profile.
- Third-party tools such as Harzing’s Publish or Perish retrieve citation data (commonly from Google Scholar) and compute h-index and related variants without requiring a saved profile; useful for a one-off check or for a database that doesn’t offer a persistent author profile.
Calculation edge cases worth knowing about
- Self-citations are included by default. None of the three databases excludes a researcher’s citations of their own prior work from the base h-index calculation shown on a profile; some platforms offer a separate “excluding self-citations” view as an optional filter rather than changing the default number.
- Co-authored papers count in full toward every co-author’s individual h-index — the h-index does not divide credit by author count or authorship position, which is one reason it can be a poor comparator across fields with very different typical team sizes.
- Name and profile disambiguation matters more than the arithmetic. An incomplete or duplicate author profile (common with common surnames, name changes, or institutional-affiliation variants) will undercount publications before the h-index calculation even starts. A persistent identifier such as an ORCID iD, where linked into a database profile, reduces this risk.
- The h-index can only stay the same or increase over time within a fixed database and pull date — new citations can only add support to already-qualifying papers or push a paper past the next rank threshold, never remove a citation already counted. A displayed h-index can appear to drop only if the underlying database changes its indexed content (e.g., a paper or citing source is delisted) or if it’s compared across different pull dates using different coverage.
Which coverage difference is producing your gap
Knowing that the three numbers differ is not the same as knowing why yours differ, and a promotion committee that asks “why is your Google Scholar number so much higher?” is asking the second question. The gap for any individual researcher is the sum of a small number of identifiable, checkable mechanisms. Each one below is stated from the database’s own published policy, with the check you can run against your own record.
1. Theses, dissertations and technical reports
Google Scholar’s Inclusion Guidelines for Webmasters name the eligible document types directly: “journal papers, conference papers, technical reports, or their drafts, dissertations, pre-prints, post-prints, or abstracts.” A doctoral thesis deposited in an institutional repository is therefore a first-class citing document in Scholar. Scopus’s content policy and selection criteria are built around serial titles with an ISSN and a documented peer-review process, so a repository-deposited thesis is not a Scopus source and its citations to you do not exist inside Scopus at all. Check: in your Scholar profile, open a paper’s “Cited by” list and count how many citing items are theses or university repository records. In a supervision-heavy field that single category can account for most of the gap on its own.
2. Preprints, and the year the preprint clock starts
Scholar indexes “pre-prints, post-prints” as eligible content with no start date. Scopus states that its preprint coverage runs from 2017 onward. Web of Science handles preprints through a separate Preprint Citation Index rather than the Core Collection that its author metrics are computed from. The practical effect is asymmetric by career stage: a researcher whose early work was cited mainly in preprints before 2017 carries a permanent, structural gap that will never close, while a current-day preprint citation may be counted by two of the three.
3. Books, book chapters, and the 5MB rule
This is the mechanism almost nobody knows about. Scholar’s crawl guidelines state that “each file must not exceed 5MB in size” and that larger documents — explicitly “books and long dissertations” — must instead be uploaded to Google Book Search, from which Scholar then draws. That is not a coverage guarantee: it is a dependency on whether your publisher put the book into Google Books at all. Scopus and Web of Science both index books and book series, but each through a separate curated programme with its own selection process, not automatically. Check: if you work in a monograph-heavy discipline, list your chapters and look up each one individually in all three; a chapter that returns zero citations in one database is usually absent as a source, not uncited.
4. Non-English and regional journals
Scopus’s technical eligibility requires that a title “have content that is relevant for and readable by an international audience (have English language abstracts and titles).” A regional journal publishing in Spanish, Portuguese, Bahasa, Turkish or Chinese without English titles and abstracts is not eligible on that criterion alone, irrespective of its scholarly quality. Google Scholar’s stated scope is the opposite: “all fields of research, all languages, all countries, and over all time periods.” For researchers whose citing community publishes substantially in a language other than English, this is typically the single largest term in the gap — and it is a coverage policy, not a quality judgment about the citing work.
5. Records that were never crawlable in the first place
Scholar’s guidelines set conditions a publisher’s site must meet before its articles are indexed at all: files must be HTML or PDF, PDFs “must have searchable text” (a scanned page image that was never OCR’d does not qualify), and the site must expose either full text or the complete author-written abstract without a login. The guidelines are explicit about the failure case: “Sites that show login pages, error pages, or bare bibliographic data without abstracts will not be considered for inclusion and may be removed from Google Scholar.” A citation published in a journal whose platform fails one of those conditions is invisible to Scholar while being perfectly visible to Scopus or Web of Science — the one direction of the gap that runs the “wrong” way, and the reason a Scholar h-index is not always the highest of the three.
6. Duplicate and split records
Because Scholar groups versions of a document algorithmically rather than against a registered identifier, one paper can appear as two entries — commonly a preprint and the version of record, or two spellings of a title. When that happens the citations split between the entries, and because the h-index depends on the citation count of individual items, a split record can hold your h-index one point lower than the underlying data supports. Merging the duplicates in your profile can change the number the same day. The reverse also occurs: an over-merged entry, or another researcher’s paper auto-added to a profile with a common name, inflates it. See troubleshooting missing or inaccurate citations on a Google Scholar profile for the merge procedure, and merging duplicate Web of Science author records for the equivalent on the Clarivate side, where a split author record does the same damage for a different reason (algorithmic name disambiguation rather than version grouping). On the Elsevier side the identity anchor is the Scopus Author ID.
7. Your own institution’s Web of Science subscription
The Web of Science h-index has a property the other two do not: it is not purely a property of your publication record. Clarivate’s Citation Report documentation states that the h-index it reports is based on the depth of years of your institution’s product subscription and your selected timespan, and that source items outside that subscription are not factored into the calculation. Two researchers with identical publication and citation records, at two institutions holding different Core Collection editions or different backfile depths, can therefore be shown different Web of Science h-index values — and the same researcher can see their own number change simply by moving institutions or by altering the timespan filter. Check: before recording a Web of Science figure, note which editions and which year range your session is actually searching, and record them alongside the number. See how to read a Web of Science citation report.
8. Citations from sources the database does not index
The mechanism underneath all of the above: each database counts a citation only when the citing document is itself inside that database. A citation to your work from a government report, a clinical guideline, a policy document, a patent or a non-indexed regional journal is real scholarly uptake that Scopus and Web of Science structurally cannot count, because the citing item is not a source in their universe. Scholar counts it if it crawled it. This is why the gap tends to be widest in applied, policy-facing and practice-facing fields, and narrowest in fields whose entire citing literature sits in indexed journals.
How to audit your own gap
The point of the audit is to be able to say, in one sentence to a committee, which mechanisms account for your difference. It takes about twenty minutes:
- Pull all three h-index values on the same day and write down the date. They drift independently; a comparison across different pull dates is not a comparison.
- Identify your h-core in each database — the set of papers at or above the h threshold. The gap almost never comes from your whole record; it comes from a handful of papers that sit just above the line in one database and just below it in another.
- Take the two or three boundary papers and compare their citing lists item by item. Categorise each citing document Scholar has but Scopus or Web of Science does not into the buckets above: thesis, preprint, book chapter, non-English journal, report, duplicate.
- Check the reverse direction too. List anything Scopus or Web of Science counts that Scholar missed — that usually points at mechanism 5 (an uncrawlable publisher platform) or at a record missing from your Scholar profile entirely.
- Fix what is fixable before you report anything. Merge split records, remove misattributed papers, and confirm your author identifiers are clean. Only the gap that survives that cleanup is a real coverage difference; the rest was a profile-hygiene problem. To build the underlying record once and reuse it, see compiling one complete publication list across PubMed, Scopus, Web of Science and Google Scholar.
What you take into the dossier is then a labelled figure plus a one-line explanation of the difference — for example, that the Scholar figure is higher because the citing literature in the field is substantially non-English and thesis-based, both of which are outside Scopus’s stated eligibility criteria. That sentence is what converts an apparent inconsistency into evidence that you understand your own bibliometrics. (See the FAQ below on which database’s number to report.)
What cannot be verified, and what to do about it
Two limits are worth stating plainly, because pages that gloss over them mislead:
- Google Scholar publishes no coverage figures and no source list. The Inclusion Guidelines describe what a publisher’s website must do to be indexed; they are not a statement of what Scholar has actually indexed, and Google has never released an index size, a title list, or a deduplication methodology. Any claim about Scholar’s total coverage is an outside estimate, not a vendor-published figure — treat it accordingly.
- Clarivate’s detailed product documentation sits behind institutional authentication and its help portal has migrated hosts, so the subscription-dependency point above should be confirmed against your own library’s Web of Science entitlement rather than assumed. Journal Citation Reports in particular is not publicly reachable; this site records it as unverifiable from the open web. Your librarian can tell you in a minute which editions and years your institution actually holds.
For how a fourth, open database changes this picture — and the coverage-versus-curation trade-off it makes explicit — see this site’s guide to OpenAlex, and the Scopus vs. Web of Science vs. OpenAlex comparison. For the two pairwise database comparisons underlying the mechanisms above, see Google Scholar vs. Scopus and Google Scholar vs. Web of Science. For the field-level equivalent of the same coverage question — citation thresholds computed inside a single curated universe — see Essential Science Indicators. And for finding the underlying records efficiently in Scholar itself, see which Google Scholar advanced search operators actually work.
Frequently asked questions
Which database’s h-index should I report?
Report the number together with the source database and the date it was pulled (for example, “h-index of 14, Scopus, as of March 2026”) rather than a bare number. Institutions and funders that request an h-index in a CV or application will often specify which database they expect; when they don’t, naming the source avoids the figure being misread as directly comparable to a colleague’s number from a different database.
Is a higher h-index always better?
Not straightforwardly. The h-index is field-dependent (citation density varies enormously by discipline), career-stage-dependent (it can only grow, so it favors longer careers), and blind to authorship order or contribution. See what counts as a “good” h-index for how field and career stage change what a given number means, and DORA/CoARA guidance on why it should not be used as a standalone decision criterion.
Does the h-index count preprints?
Only if the database computing it indexes preprints. Google Scholar generally does; Scopus and Web of Science’s core indexing has historically focused on peer-reviewed content, though both have expanded preprint and early-access coverage over time — check the specific database’s current scope rather than assuming.
How is h-index different from i10-index?
The i10-index is a simpler count — the number of a researcher’s papers with at least 10 citations each — and is a Google Scholar–specific metric that Scopus and Web of Science don’t compute. See the h-index vs. i10-index comparison for the full breakdown, and impact factor vs. h-index for how an author-level metric differs from a journal-level one.
This guide covers calculation mechanics and database differences. For what the number means once you have it — typical ranges by field and career stage, and why assessment frameworks caution against using it as a standalone metric — see the h-index and good h-index dictionary entries.








