Written and maintained by CASRAI Editorial Board
Last updated
Last verified: August 31, 2026, against the AGREE Next Steps Consortium’s own AGREE II User’s Manual and Brouwers MC, Kho ME, Browman GP, et al. “AGREE II: Advancing guideline development, reporting and evaluation in health care,” CMAJ 2010;182(18):E839-E842, and its companion validity paper, “Development of the AGREE II, part 2,” CMAJ 2010;182(10):E472-E478. This is a methodological orientation, not a substitute for the full instrument and user’s manual when actually conducting an appraisal.
What AGREE II is, and what it isn’t
AGREE II (Appraisal of Guidelines for Research & Evaluation II) is a validated instrument for judging the methodological quality of a clinical practice guideline — the document a panel produces when it turns a body of evidence into a specific recommendation (“use drug X for condition Y”). It was developed by the AGREE Next Steps Consortium, an international collaboration of guideline developers and methodologists, and published in 2010 as an update to the original 2003 AGREE Instrument. The current instrument and manual carry a 2017 reprint date but reflect no substantive change to the 23 items themselves.
The distinction worth getting right before using it: AGREE II does not appraise a clinical trial, a systematic review, or a diagnostic-accuracy study. It appraises the guideline — the synthesis-and-recommendation document that typically sits downstream of that primary evidence, often produced by a GRADE panel. A systematic review feeding into a guideline gets appraised with a tool like AMSTAR 2; the guideline itself gets appraised with AGREE II. Confusing the two is common, because both produce a checklist-shaped score and both sit in the evidence-synthesis space, but they are scoring different objects.
The six AGREE II domains and 23 items
AGREE II organizes 23 items into six independent domains, each scored on a 7-point scale (1 = strongly disagree, 7 = strongly agree) by each appraiser, per item:
- Scope and Purpose (items 1–3) — whether the guideline’s overall objective, the health question(s) it covers, and the population it applies to are specifically described.
- Stakeholder Involvement (items 4–6) — whether the guideline development group includes the right professional groups, whether patient/public views and preferences were sought, and whether target users are clearly defined.
- Rigour of Development (items 7–14) — the largest domain: systematic methods for searching evidence, explicit criteria for selecting it, clear description of the evidence’s strengths and limitations, explicit methods for formulating recommendations, consideration of health benefits/harms/risks, an explicit link between recommendations and the supporting evidence, external review before publication, and a procedure for updating the guideline.
- Clarity of Presentation (items 15–17) — whether recommendations are specific and unambiguous, whether different management options for the condition are clearly presented, and whether key recommendations are easily identifiable.
- Applicability (items 18–21) — facilitators/barriers to implementation, advice or tools for putting recommendations into practice, consideration of resource implications, and monitoring/auditing criteria.
- Editorial Independence (items 22–23) — whether the views of the funding body did not influence the guideline’s content, and whether competing interests of development group members were recorded and addressed.
Two further items sit outside the six domains as a separate Overall Assessment: a global 1–7 rating of the guideline’s overall quality, and a categorical judgment on whether the appraiser would recommend the guideline for use (yes / yes, with modifications / no).
How AGREE II scoring works
Each domain is scored, not the instrument as a whole — the AGREE Next Steps Consortium is explicit that the six domain scores are independent and should not be summed or averaged into one overall number. A domain’s scaled score is calculated as:
(obtained score − minimum possible score) ÷ (maximum possible score − minimum possible score) × 100
where the minimum and maximum possible scores are set by the number of items in that domain, the number of appraisers, and the 7-point scale. A guideline can score well on Scope and Purpose while scoring poorly on Editorial Independence, and AGREE II is designed to surface exactly that kind of unevenness rather than average it away. The instrument’s manual recommends using at least two appraisers, and ideally four, per guideline — a single appraiser’s ratings are far less reliable, and AGREE II’s own validation work was built on multi-rater agreement, not solo scoring.
Who uses AGREE II, and when
Three overlapping audiences apply it, for different reasons:
- Guideline developers use it prospectively, during development, as a quality checklist — catching a missing conflict-of-interest disclosure or an undocumented evidence-selection process before publication rather than after.
- Guideline users — a hospital committee, a specialty society, an individual clinician — use it to decide whether to adopt or endorse an existing guideline, particularly when two guidelines from different bodies disagree.
- Researchers and methodologists use it comparatively, appraising multiple guidelines on the same clinical question as part of a guideline-quality review or an overview of guidelines.
The free online tool My AGREE PLUS, hosted by the AGREE Enterprise, lets appraisers score a guideline against the instrument and generates the domain scores automatically rather than requiring manual calculation.
AGREE II vs. reporting guidelines: appraising a guideline is not the same job as reporting a study
This is the distinction most worth stating plainly, because CASRAI’s own reporting-guideline content covers adjacent but structurally different ground. A reporting guideline — CONSORT for randomized trials, PRISMA for systematic reviews, STARD for diagnostic-accuracy studies, CONSORT-AI/TRIPOD+AI for AI-involved studies — tells an author what to disclose so a reader can judge whether a specific study was conducted and reported transparently. AGREE II asks a different question of a different object: not “did this guideline’s authors report their process transparently,” but “was the process itself methodologically sound” — did the group search the evidence systematically, weigh the tradeoffs, and manage its conflicts of interest, regardless of how tidily that gets written up.
The AGREE Next Steps Consortium maintains a companion AGREE Reporting Checklist specifically to close that gap for guideline developers — it plays the reporting-guideline role that CONSORT/PRISMA/STARD play for trials, reviews, and diagnostic studies, but scoped to guideline documents. AGREE II itself remains the appraisal instrument, used by a reader or adopter of the guideline, not the reporting checklist used by its authors. A CASRAI reader working across this territory is usually reaching for one of three tools: a critical appraisal tool (AGREE II, AMSTAR 2, the study-design-matched options that guide covers) to judge quality after the fact, a reporting guideline to write a transparent report in the first place, or GRADE to grade the certainty of the underlying evidence base a guideline’s recommendations rest on. All three interlock on a real guideline project; none substitutes for the other two.
The AGREE family: related tools
- AGREE-REX (Recommendation EXcellence) — a companion instrument focused specifically on the quality of individual recommendations, rather than the guideline development process as a whole.
- AGREE-HS (Health Systems) — adapted for appraising guidance aimed at health-system or policy-level decisions rather than individual clinical care.
- AGREE Reporting Checklist — the reporting-guideline counterpart described above, for guideline developers writing up their process.
- My AGREE PLUS — the free online appraisal and scoring tool hosted by the AGREE Enterprise.
Frequently asked questions
What is AGREE II used for?
Judging the methodological quality and transparency of a clinical practice guideline’s development process — not the underlying clinical evidence itself, and not a single study or systematic review.
How many items does AGREE II have?
23 items across six domains (Scope and Purpose, Stakeholder Involvement, Rigour of Development, Clarity of Presentation, Applicability, Editorial Independence), plus two separate Overall Assessment items.
Who developed AGREE II?
The AGREE Next Steps Consortium, building on the original 2003 AGREE Instrument; the validated update was published in CMAJ in 2010.
Is AGREE II free to use?
Yes — the instrument, user’s manual, and the My AGREE PLUS online scoring tool are freely available from the AGREE Enterprise.
How is an AGREE II domain score calculated?
(Obtained score − minimum possible score) ÷ (maximum possible score − minimum possible score) × 100, calculated separately for each of the six domains. Domain scores are never combined into one overall number.
What’s the difference between AGREE II and PRISMA or CONSORT?
PRISMA and CONSORT are reporting guidelines: they tell an author what to disclose when writing up a systematic review or a randomized trial. AGREE II is an appraisal instrument: it judges the methodological rigor of a clinical practice guideline, a different kind of document produced through a different process.








