Written and maintained by CASRAI Editorial Board
Last updated
What the Downs and Black Checklist Is
The Downs and Black checklist is a critical-appraisal tool published by Sue Downs and Nick Black in 1998 in the Journal of Epidemiology and Community Health, under the title “The feasibility of creating a checklist for the assessment of the methodological quality both of randomised and non-randomised studies of health care interventions.” That title is also the point of the tool: rather than building separate instruments for trials and observational studies, the authors set out to test whether a single checklist, grounded in standard epidemiological principles, could assess both study designs on comparable terms.
The checklist was developed and piloted against a set of published health-intervention studies, then tested for internal consistency, test-retest reliability and inter-rater reliability. The original paper reports high internal consistency for the overall Quality Index (Kuder-Richardson KR-20 coefficient of 0.89) and solid test-retest reliability, with an assessment taking roughly 20 minutes per study once a reviewer is familiar with the items. The authors’ own conclusion was measured rather than promotional: it is feasible to build a checklist usable across both randomised and non-randomised designs, though they flagged that the external-validity items in particular needed further refinement — a caveat worth keeping in mind when the tool is applied uncritically today.
Structure: 27 Items Across Five Sections
The instrument is organised into five sections that map onto the standard quality domains used in study appraisal: how the study is reported, whether its findings generalise (external validity), whether its internal comparisons are free of bias, whether they are free of confounding, and whether the study had adequate statistical power to detect the effect it was looking for.
- Reporting (10 items) — whether the hypothesis/aim, main outcomes, patient characteristics, interventions, confounders, main findings, estimates of random variability, adverse events, losses to follow-up, and actual probability values are clearly described.
- External validity (3 items) — whether the subjects recruited, and the setting/facilities they were recruited into, are representative of the population the study’s conclusions are drawn about, and whether those who agreed to participate were representative of the population from which they were recruited.
- Internal validity — bias (7 items) — blinding of outcome assessors and subjects, appropriateness of statistical tests, reliability of the main outcome measures, compliance with the intervention, and whether the time periods being compared were contemporaneous.
- Internal validity — confounding and selection bias (6 items) — whether subjects were recruited from the same population and over the same time period, randomised, whether allocation was concealed, adequacy of adjustment for confounding in the analysis, and whether losses to follow-up were accounted for.
- Power (1 item) — whether the study had sufficient statistical power to detect a clinically important effect, scored against a lookup table of sample sizes and effect sizes rather than a simple yes/no.
That gives 27 items in total. Scoring is not uniform across items: most are scored 0 or 1, one reporting item is scored 0-2, and the power item is scored on a 0-5 scale via the lookup table, which is why the checklist’s maximum achievable score is commonly cited as running into the low 30s rather than a flat 27. Reviewers using the tool should reproduce the original scoring table rather than assume every item is worth one point — a frequent, avoidable source of miscalculation.
Why It’s One of the Few Tools Usable Across Both Study Designs
Most widely-used critical-appraisal instruments are design-specific. RoB 2 assesses randomised trials; ROBINS-I and ROBINS-E assess non-randomised studies of interventions and exposures respectively. That split is defensible on methodological grounds — randomisation removes a category of confounding that no amount of statistical adjustment fully replaces — but it creates a practical problem for a systematic review that deliberately mixes RCTs and observational studies in the same synthesis, which is common when trial evidence on an intervention is sparse and cohort or case-control data has to fill the gap.
The Downs and Black checklist’s cross-design applicability is exactly what makes it useful there: applying the same instrument to every included study, whatever its design, gives a single comparable quality score or profile across the whole evidence base, rather than forcing a review team to report two incommensurable sets of appraisal results side by side. This is the main reason it persisted in the evidence-synthesis literature well after design-specific tools like the Cochrane risk-of-bias tools became standard for single-design reviews — see choosing a review type for where a mixed-design synthesis is the right call in the first place, and where it isn’t.
Practical Limitations
Two limitations are consistently raised in the methodology literature and are worth stating plainly rather than glossing over, given the “quality bar” this page is trying to meet:
- Scoring complexity. The non-uniform item weights (0-1, 0-2, 0-5) and the separate lookup table required for the power item make the checklist more laborious to apply consistently than a flat yes/no tool, and many reviews using it either simplify the power item to a binary score or drop it altogether — a deviation from the original instrument that should be disclosed in the review’s methods section, not left implicit.
- Inter-rater variability. Because several items (adequacy of confounder adjustment, representativeness of the recruited sample, whether losses to follow-up were adequately described) call for a judgement rather than a simple lookup, independent reviewers applying the checklist to the same study can reach different scores on those items even after training — the same class of reviewer-judgement variability that motivated the more structured, domain-based signalling questions in later tools like ROBINS-I and RoB 2. In practice this means two raters per study and a documented disagreement-resolution process are not optional extras; they’re what keeps the checklist’s scores comparable across a review team, and the original paper’s own emphasis on reliability testing is a signal that the authors expected this to be checked, not assumed.
The authors’ own paper also flagged the external-validity section specifically as needing further work, and later users have echoed that the three external-validity items ask reviewers to judge “representativeness” against a population that the source studies rarely define with enough precision to score confidently — a limitation inherited from the checklist’s era rather than fixed by later revisions.
Downs and Black vs Other Critical-Appraisal Tools
Where a review is entirely made up of one study design, a design-specific tool is usually the better default: RoB 2 for RCTs, or ROBINS-I for non-randomised intervention studies. Generic checklists that, like Downs and Black, aim for broader applicability include the CASP checklists (design-specific but easier to apply, without a unified cross-design score) and the JBI critical appraisal checklists, which cover an even wider range of designs including qualitative and prevalence studies. Where a review deliberately combines quantitative and qualitative evidence, the Mixed Methods Appraisal Tool (MMAT) is the more current purpose-built option rather than stretching Downs and Black beyond intervention studies. None of these substitute for reporting checklists like PRISMA, which governs how the review itself is written up, not how the quality of its included studies is judged — Downs and Black sits alongside, not in place of, that reporting layer.
When to Use It in a Systematic Review
Downs and Black is the right tool when a review’s inclusion criteria genuinely span randomised and non-randomised intervention studies and the team wants one consistent scoring framework rather than two. It is a weaker choice when a review is design-homogeneous (use the matching design-specific tool instead), when the analysis plan calls for domain-level risk-of-bias judgements rather than a summary score (the Cochrane RoB 2 / ROBINS-I traffic-light approach is built for that), or when the review will feed into a GRADE certainty-of-evidence assessment, which expects domain-based risk-of-bias ratings as an input rather than a single composite score. Whichever tool is chosen, the methods section should record it, the version/modifications used, the number of independent raters, and how disagreements were resolved — the same transparency the checklist’s own reliability testing was designed to support.
Frequently Asked Questions
Is the Downs and Black checklist still recommended for new systematic reviews?
It remains a legitimate option, particularly for reviews mixing randomised and non-randomised intervention studies, but the Cochrane Handbook’s current guidance favours design-specific tools such as RoB 2 and ROBINS-I for reviews within a single design, precisely because of the domain-level detail those tools provide over a single composite score. Where a review genuinely needs one instrument across mixed designs, Downs and Black is still cited and used; where it doesn’t, a design-specific tool is usually the better fit.
How many items does the Downs and Black checklist have, and what is the maximum score?
27 items across five sections: reporting (10), external validity (3), internal validity — bias (7), internal validity — confounding (6), and power (1). Most items score 0 or 1, one reporting item scores 0-2, and the power item is scored 0-5 against a lookup table, so the achievable maximum is higher than 27 — reviewers should use the original scoring table rather than assume a flat point-per-item scheme.
Can Downs and Black be used for a single-design review of only RCTs?
It can, but there is little reason to: RoB 2 was purpose-built for randomised trials and gives domain-level bias judgements that a composite Downs and Black score does not. Downs and Black’s comparative advantage is specifically the ability to score randomised and non-randomised studies on the same scale within one review.
What is the biggest source of disagreement between raters using this checklist?
Items requiring a judgement call rather than a factual lookup — adequacy of confounder adjustment, representativeness of the study sample, and completeness of loss-to-follow-up reporting — are where independent raters most often diverge, which is why using two raters with a documented adjudication process is standard practice rather than an optional refinement.
Primary source: Downs SH, Black N. “The feasibility of creating a checklist for the assessment of the methodological quality both of randomised and non-randomised studies of health care interventions.” J Epidemiol Community Health. 1998;52(6):377-384.








