Written and maintained by CASRAI Editorial Board
Last updated
ROBIS (Risk Of Bias In Systematic reviews) is a tool for assessing whether a systematic review itself was conducted with enough rigor to trust its conclusions — as distinct from tools that assess the risk of bias in the individual studies a review includes, such as RoB 2 or ROBINS-I. It was developed by Penny Whiting and colleagues at the University of Bristol’s School of Social and Community Medicine, working with international collaborators as the ROBIS group, and published in the Journal of Clinical Epidemiology in 2016 (Whiting P, Savović J, Higgins JPT, et al. “ROBIS: A new tool to assess risk of bias in systematic reviews was developed.” J Clin Epidemiol. 2016;69:225–234. doi: 10.1016/j.jclinepi.2015.06.005).
Last verified 2026-08-31 against the original 2016 J Clin Epidemiol publication and the ROBIS project’s own guidance document (University of Bristol).
The Three Phases of ROBIS
ROBIS is completed as a sequence of three phases, applied by someone assessing an already-completed systematic review:
- Phase 1 — Assess relevance (optional). A quick check of whether the review’s scope actually matches the assessor’s research question, useful when ROBIS is being used to screen many reviews at once (for example, in an overview of reviews).
- Phase 2 — Identify concerns with the review process. The main body of the assessment, working through four domains of the review’s conduct (below) and answering a set of signalling questions under each.
- Phase 3 — Judge risk of bias. Using the domain-level judgments from Phase 2, the assessor reaches an overall risk-of-bias judgment for the review’s own conclusions.
Phase 2’s Four Domains
Bias can enter a systematic review at several distinct points in how it was conducted, and ROBIS structures Phase 2 around four of them:
- Study eligibility criteria. Whether the review’s inclusion/exclusion criteria, and the populations, interventions, comparators and outcomes they define, were appropriate and pre-specified rather than adjusted after seeing the results.
- Identification and selection of studies. Whether the search was comprehensive enough (databases, grey literature, date/language restrictions) and whether study selection and any risk-of-bias screening minimized error and bias.
- Data collection and study appraisal. Whether data extraction and the appraisal of included studies’ own risk of bias were done in a way that minimized error, ideally in duplicate.
- Synthesis and findings. Whether the methods used to combine or synthesize the included studies’ results were appropriate given their design and heterogeneity, and whether all pre-specified outcomes and analyses were reported.
Each domain is worked through with a set of signalling questions (answered Yes, Probably Yes, Probably No, No, or No Information), which then support a domain-level risk-of-bias judgment before the assessor moves to Phase 3.
Phase 3: The Overall Judgment
Phase 3 asks the assessor to weigh the four domain-level judgments together and reach one overall risk-of-bias rating for the review’s conclusions:
| Rating | What it means |
|---|---|
| Low | The review’s conduct across all four domains was unlikely to have introduced a meaningful bias into its conclusions. |
| High | At least one domain raises concerns serious enough that they could have influenced the review’s conclusions. |
| Unclear | There isn’t enough reported detail about the review’s conduct in one or more domains to make a confident judgment either way. |
This final judgment is about the review’s conclusions, not a mechanical tally of “Yes” answers — a review can have a minor limitation in one domain and still be judged low risk of bias overall if that limitation was unlikely to change its findings.
ROBIS vs AMSTAR 2
ROBIS and AMSTAR 2 are both applied to a completed systematic review rather than to the studies inside it, and they’re often confused for the same job, but they answer different questions:
- ROBIS asks specifically whether the review’s conduct introduced bias into its conclusions — it was designed as, in the original authors’ framing, the first tool built specifically for that narrower question.
- AMSTAR 2 asks the broader question of the review’s overall methodological quality, across 16 items covering things ROBIS doesn’t score directly, such as whether the protocol was registered in advance or whether the review disclosed funding sources and conflicts of interest.
Guideline developers and authors of overviews of reviews sometimes apply both: AMSTAR 2 as a broader quality screen, ROBIS where the specific question is whether a given review’s conclusions can be trusted despite how it was conducted. Neither tool appraises the primary studies included in the review — that’s the job of study-level tools like RoB 2, ROBINS-I, ROBINS-E, or the Newcastle-Ottawa Scale, depending on study design — see choosing a critical appraisal tool for how all of these fit together by design and use case.
Who Uses ROBIS, and When
The ROBIS authors describe its primary audience as guideline developers, authors of overviews of systematic reviews (“reviews of reviews”), and review authors themselves who want to check or avoid risk of bias in their own review before submission. It’s currently scoped to four broad review types, mainly within healthcare: intervention reviews, diagnostic test accuracy reviews, prognosis reviews, and etiology/risk-factor reviews. A reader building or appraising a systematic review under PRISMA, registering one on PROSPERO, or peer reviewing a systematic review submitted to a journal is the typical person reaching for ROBIS.
Frequently Asked Questions
Is ROBIS the same as a risk-of-bias assessment of the individual studies in a review?
No. ROBIS assesses the systematic review itself — how it was conducted — not the individual primary studies it includes. Those are assessed separately with a study-level tool appropriate to their design, such as RoB 2 for randomized trials or ROBINS-I/ROBINS-E for non-randomized studies.
Do I need to complete Phase 1 of ROBIS?
No, Phase 1 (assessing relevance) is explicitly optional in the original tool — it exists to help someone screening many reviews at once (for example, for an overview of reviews) quickly set aside reviews that don’t match their research question before spending time on the full Phase 2 assessment.
Can a systematic review be judged low risk of bias by ROBIS but still be flagged by AMSTAR 2?
Yes. The two tools score different things, so their outputs are not interchangeable. A review could be conducted rigorously enough that ROBIS finds low risk of bias in its conclusions, while still scoring lower on AMSTAR 2 because of a broader methodological-quality item ROBIS doesn’t assess, such as an unregistered protocol.








