Written and maintained by CASRAI Editorial Board
Last updated
QUADAS-2 is the standard tool for assessing risk of bias and applicability in primary studies included in a systematic review of diagnostic test accuracy. It was published by Whiting and colleagues in the Annals of Internal Medicine in 2011 as a revision of the original QUADAS tool, and it is now the tool recommended by the Cochrane Collaboration for diagnostic test accuracy (DTA) reviews. This guide covers the four domains QUADAS-2 assesses, how its signalling questions are used to reach a risk-of-bias judgment, how the tool differs from QUADAS, and how it differs from risk-of-bias tools built for other study designs.
What QUADAS-2 Assesses
A diagnostic accuracy study compares an index test — the test whose accuracy is being evaluated — against a reference standard, the best available method for confirming whether the target condition is truly present or absent. QUADAS-2 is built specifically around that comparison: it evaluates whether the way patients were selected, the way the index test was performed and interpreted, the way the reference standard was performed and interpreted, and the timing/completeness of testing could have distorted the reported sensitivity and specificity. It does not evaluate treatment effect, and it has no domain for randomisation or comparator groups, because a diagnostic accuracy study has neither — every enrolled patient receives both the index test and the reference standard (or is meant to).
QUADAS-2 assesses two separate things for most domains: risk of bias (could the way the study was conducted have distorted the accuracy estimate?) and applicability (does the study’s patient population, index test, or reference standard actually match the review question the systematic review is trying to answer?). A study can be at low risk of bias and still be a poor match for the review’s applicability — for example, a test validated in a specialist referral population when the review question concerns primary-care screening.
The Four Domains
QUADAS-2 comprises four domains, each anchored to a specific part of the study’s design and conduct. The first three are assessed for both risk of bias and applicability; the fourth (flow and timing) is assessed for risk of bias only, since patient flow and test timing have no independent applicability dimension of their own.
1. Patient Selection
This domain asks whether the method of selecting participants could have introduced bias, and whether the included patients match the review question. Signalling questions used to inform the judgment:
- Was a consecutive or random sample of patients enrolled?
- Was a case-control design avoided?
- Did the study avoid inappropriate exclusions?
A case-control design — recruiting confirmed cases separately from confirmed non-cases, rather than testing a consecutive clinical population — tends to inflate sensitivity and specificity because the two groups are often more clearly distinguishable than patients seen in routine practice. Exclusions that remove diagnostically difficult or borderline patients have the same inflating effect. For applicability, the reviewer asks whether the patients, setting, and disease spectrum in the study reflect the population the review is meant to inform.
2. Index Test
This domain covers how the test under evaluation was performed and interpreted. Signalling questions:
- Were the index test results interpreted without knowledge of the results of the reference standard?
- If a threshold was used, was it pre-specified?
Interpreting the index test with knowledge of the reference-standard result (or vice versa, in the next domain) is a form of review bias that can inflate agreement between the two. A threshold chosen after seeing the data (a post-hoc, “optimal” cut-point) will fit that dataset better than it will perform on a new one, which inflates reported accuracy. Applicability here asks whether the index test, and how it was conducted or interpreted, matches how it would actually be used in the review’s target setting — a different device, different reader expertise, or a different threshold than the one the review intends to evaluate would all raise applicability concerns.
3. Reference Standard
This domain assesses whether the reference standard used is capable of correctly classifying the target condition, and whether it was interpreted blind to the index test. Signalling questions:
- Is the reference standard likely to correctly classify the target condition?
- Were the reference standard results interpreted without knowledge of the results of the index test?
Not every “gold standard” is equally accurate, and a review team has to make an explicit judgment, tailored to the condition under review, about whether the reference standard used in each included study is fit for that purpose. Applicability concerns arise when the reference standard defines the target condition differently than the review question does — for example, a purely clinical diagnosis used as the reference standard in a review that wants accuracy against a histopathological definition.
4. Flow and Timing
This domain covers what happened to patients after enrolment: whether everyone got the same reference standard, whether the interval between index test and reference standard was short enough that the target condition’s status could not have changed in between, and whether everyone who was tested was included in the final analysis. Signalling questions:
- Was there an appropriate interval between the index test and the reference standard?
- Did all patients receive the same reference standard?
- Were all patients included in the analysis?
Excluding patients with indeterminate or inconclusive index-test results, or applying a different (often less rigorous) reference standard to patients who tested negative, are both well-documented sources of bias in diagnostic accuracy studies — the latter is sometimes called partial verification or differential verification bias. This domain has no separate applicability rating, because patient flow does not change what population, test, or reference standard the study represents; it only affects how trustworthy the reported numbers are.
How Signalling Questions Lead to a Risk-of-Bias Judgment
Each signalling question is answered yes, no, or unclear, based only on what is reported in the study (not on what the reviewer assumes was probably done). A “yes” answer to every signalling question in a domain supports, but does not automatically produce, a judgment of low risk of bias for that domain. If any signalling question is answered “no,” the domain is at higher risk of bias unless the reviewer has a specific, documented reason to conclude otherwise. An “unclear” answer means the study did not report enough detail to judge, and the domain is typically rated unclear risk of bias rather than assumed low risk — QUADAS-2 explicitly discourages treating unreported information as if it were favourable.
Signalling questions are a structured aid to judgment, not a scoring algorithm: QUADAS-2 does not sum signalling-question answers into a numeric score, and the original authors are explicit that it should never be used that way. The final call for each domain — low, high, or unclear risk of bias, and (for the first three domains) low, high, or unclear concern regarding applicability — is a reviewer judgment informed by the signalling questions, documented with the reasoning behind it.
The Four Phases of Applying QUADAS-2
Whiting et al. describe QUADAS-2 as applied in four phases, and the tool is explicitly meant to be tailored to each review rather than used identically across every topic:
- Summarise the review question — state the target condition, index test, reference standard, and patient population the review is asking about.
- Tailor the tool and produce review-specific guidance — adapt the generic signalling questions and write guidance on what counts as, say, an “appropriate interval” for this specific target condition, before appraisal starts.
- Construct a flow diagram for each primary study, showing how many patients were enrolled, tested with the index test and reference standard, and included in the final analysis — this is what the flow-and-timing signalling questions are actually answered against.
- Judge bias and applicability for each domain, using the tailored signalling questions and the flow diagram.
This tailoring step is one of the more commonly skipped parts of using QUADAS-2 in practice — teams often apply the generic wording unmodified. Review-specific guidance matters because “appropriate interval” or “correctly classify the target condition” mean something different for, say, a rapid antigen test for an acute infection than for an imaging test for a slowly progressive condition.
QUADAS-2 vs the Original QUADAS Tool
QUADAS-2 replaced the original QUADAS tool, published in 2003, after several years of use exposed practical problems. The main changes:
- Structure. The original QUADAS was a flat 14-item checklist with no domain grouping. QUADAS-2 organises items into the four domains above, each mapped to a specific stage of the study.
- Signalling questions. QUADAS-2 introduced signalling questions as an explicit aid to reaching each domain-level judgment; the original tool asked reviewers to rate each of the 14 items directly as yes/no/unclear with no equivalent structured prompt.
- Separation of bias and applicability. QUADAS-2 splits risk-of-bias judgments from applicability judgments explicitly, for each of the first three domains. The original QUADAS blended quality and applicability concerns within the same 14 items, which made it harder to tell whether a study was flagged for being poorly conducted or simply for not matching the review’s population.
- Removal of scoring. Some review teams had used QUADAS’s 14 items to generate a summary quality score. QUADAS-2 does not support scoring at all — the four-domain, judgment-based format was designed specifically to discourage that practice.
- Tailoring built in. QUADAS-2’s four-phase process (above) makes review-specific customisation an explicit step; the original tool was more often applied as a fixed, generic checklist.
A related, later extension is worth flagging for reviews that compare two or more index tests head-to-head: QUADAS-C (2021) adapts the same four-domain structure for comparative diagnostic accuracy studies, adding signalling questions about whether the comparison between tests was itself fair (for example, whether both tests were interpreted by the same readers under the same conditions).
QUADAS-2 vs Risk-of-Bias Tools for Other Study Types
QUADAS-2 is specific to diagnostic test accuracy studies. It is not a general-purpose risk-of-bias tool, and using it (or a tool from another family) on the wrong study design produces a mismatch that a peer reviewer will catch:
| Study design | What is being evaluated | Tool |
|---|---|---|
| Diagnostic test accuracy (index test vs. reference standard) | Sensitivity/specificity distortion, applicability of test and population | QUADAS-2 |
| Randomised controlled trial of an intervention | Treatment effect distortion from randomisation, blinding, attrition, reporting | Cochrane RoB 2 |
| Non-randomised study of an intervention | Confounding, selection into intervention groups, and the same domains RoB 2 covers where relevant | ROBINS-I |
| Systematic review as a whole (not its included studies) | Review conduct: protocol, search comprehensiveness, screening, synthesis methods | ROBIS, AMSTAR 2 |
The distinguishing feature is what QUADAS-2 has no domain for: it does not ask about randomisation, allocation concealment, or a comparator arm, because a diagnostic accuracy study is not testing whether an intervention changes an outcome — it is testing whether a test’s result agrees with a reference standard in the same patients. Conversely, RoB 2 and ROBINS-I have no domain for a “reference standard” at all, because an effectiveness study has no equivalent concept; its comparison is between two groups of patients rather than between two measurements of the same patient. Reviewers appraising a mixed-design body of evidence (common in test-and-treat pathways or health-technology assessments) frequently need more than one tool, one per study design, applied separately rather than blended. For a broader map of which tool fits which design, see choosing a critical appraisal tool for your study design.
Common Pitfalls When Applying QUADAS-2
- Using the generic wording unmodified. Skipping phase 2 (tailoring) leaves signalling questions like “appropriate interval” undefined for the specific target condition, which produces inconsistent judgments between two reviewers appraising the same study.
- Scoring it. Summing yes/no answers into a numeric quality score is explicitly against the tool’s design and was one of the problems QUADAS-2 was built to stop.
- Assuming “unclear” means “probably fine.” An unreported detail should be rated unclear risk of bias, not defaulted to low, even when the rest of the study reads as well-conducted.
- Skipping the flow diagram. The flow-and-timing domain is difficult to judge accurately without literally drawing how many patients entered, were tested, and were analysed — discrepancies that are easy to miss in prose are usually obvious once mapped.
- Treating risk of bias and applicability as one rating. A well-conducted study in the wrong population is a real, distinct finding from a poorly conducted study — conflating the two loses information a GRADE certainty-of-evidence assessment downstream will need.
Frequently Asked Questions
How many domains does QUADAS-2 have?
Four: patient selection, index test, reference standard, and flow and timing. The first three are rated for both risk of bias and concerns regarding applicability; flow and timing is rated for risk of bias only.
Are signalling questions scored?
No. Signalling questions (answered yes/no/unclear) inform each domain’s judgment, but QUADAS-2 does not sum them into a numeric quality score — the tool’s authors explicitly designed it to discourage that practice, unlike some earlier checklist-style appraisal tools.
What is the difference between QUADAS and QUADAS-2?
QUADAS (2003) was a flat 14-item checklist with no domain structure and no signalling questions. QUADAS-2 (2011) reorganised appraisal into four domains, added signalling questions to guide each judgment, separated risk-of-bias from applicability ratings, and built review-specific tailoring into the process.
Can QUADAS-2 be used for a systematic review of a treatment’s effectiveness?
No. QUADAS-2 is specific to diagnostic test accuracy studies (index test vs. reference standard). A systematic review of an intervention’s effectiveness should use Cochrane RoB 2 for randomised trials or ROBINS-I for non-randomised intervention studies.
What is QUADAS-C?
QUADAS-C (2021) is a later extension of QUADAS-2 built for comparative diagnostic accuracy studies — reviews that assess two or more index tests head-to-head rather than one index test against a single reference standard alone. It keeps the same four-domain structure and adds signalling questions about whether the comparison between tests was conducted fairly.
Where should QUADAS-2 sit in a systematic review’s methodology?
After study selection and data extraction, alongside or shortly before synthesis. Its domain-level judgments feed directly into a diagnostic test accuracy synthesis and into any downstream GRADE certainty-of-evidence rating, the same way RoB 2 judgments feed a GRADE assessment for an intervention review — see the Cochrane Handbook chapter guide for where DTA-specific methods sit relative to the standard intervention-review chapters.








