Written and maintained by CASRAI Editorial Board
Last updated
A GRADE certainty rating tells you how confident to be in an effect estimate. It does not, by itself, tell you what to do about it. The Evidence-to-Decision (EtD) framework is the structured step in between: a standard set of criteria that a guideline panel, a health system, or a coverage-policy body works through, in the same order, to turn “the certainty of evidence is moderate” into “we recommend for” or “we suggest against, conditionally.” Developed by the GRADE Working Group and formalized in a 2016 BMJ series by Alonso-Coello and colleagues, EtD frameworks exist specifically to make that translation step transparent and reproducible instead of an unstated judgment call buried in a guideline’s discussion section.
What the EtD framework adds beyond a GRADE certainty rating
GRADE’s four-level certainty rating (high, moderate, low, very low), covered in more depth on CASRAI’s GRADE dictionary entry, answers one question: how much confidence should we place in an estimate of effect, based on risk of bias, inconsistency, indirectness, imprecision, and publication bias? That rating is an input to a recommendation, not the recommendation itself. Two guideline panels looking at the same moderate-certainty evidence can reasonably land on different recommendations if their populations weight the trade-off between benefits and harms differently, or if implementing the intervention costs one health system far more than another. The EtD framework is where that second layer of judgment happens, and it forces the panel to record each criterion’s assessment and the reasoning that connects it to the final call — the design goal being that a reader can trace exactly why the panel recommended what it did, not just what they recommended.
The criteria in a GRADE EtD framework
The original clinical-recommendation EtD framework (there are separate, related templates for coverage decisions, health-system and public-health recommendations, and diagnostic-test recommendations — see below) works through these criteria in sequence:
- Problem — is the problem a priority? Establishes whether the question is worth a panel’s time before assessing anything else.
- Desirable and undesirable effects — the magnitude of benefits and harms, drawn directly from the underlying meta-analysis or pooled estimate.
- Certainty of the evidence — the GRADE rating itself (high/moderate/low/very low), carried into the framework as one input among several rather than the sole determinant.
- Values — how much the affected population is likely to value the main outcomes, and how much variability exists in that valuation.
- Balance of effects — weighing desirable against undesirable effects once values are factored in; this is the criterion that most directly drives direction (for or against) and strength.
- Resources required and the certainty of evidence about resource use — cost and cost-effectiveness, assessed with its own separate certainty judgment because cost data often comes from weaker evidence than the clinical effect estimate.
- Equity — whether the intervention would narrow or widen health disparities between groups.
- Acceptability — whether key stakeholders (clinicians, patients, payers) would find the option acceptable in practice.
- Feasibility — whether the option can actually be implemented given real-world infrastructure and workflow constraints.
Each criterion is rated on a simple judgment scale (for example, “favors the intervention,” “favors the comparison,” “probably favors the intervention,” “no important difference,” “varies,” or “don’t know”), and the panel records the research evidence and any additional considerations behind each judgment. GRADEpro GDT (Guideline Development Tool), the software most guideline panels use to build these tables, structures the whole process this way by default, which is a large part of why the framework has become a de facto standard rather than remaining a journal-article proposal.
From criteria to recommendation strength
Once every criterion is rated, the panel makes two separate calls: direction (for or against the intervention) and strength (strong or conditional/weak). A strong recommendation means the panel is confident that the desirable consequences of an intervention clearly outweigh the undesirable ones (or vice versa, for a strong recommendation against) — most informed people would want the recommended course of action, and it is reasonable to use it as a default or performance-measure standard. A conditional (weak) recommendation means the trade-offs are closer, evidence is less certain, or values are expected to vary meaningfully across patients — the framework’s own output should prompt shared decision-making rather than a blanket default, and it is a poor basis for a rigid quality metric. This is the practical payoff of running the full EtD table rather than jumping from certainty rating straight to a recommendation: the strength label itself becomes traceable to which specific criteria pushed it toward “strong” versus “conditional.”
EtD frameworks beyond clinical recommendations
The 2016 BMJ series that formalized EtD frameworks published two linked papers, not one, because a single template doesn’t fit every decision type:
- Clinical recommendations — the version above, for guideline panels recommending an intervention to clinicians and patients.
- Coverage decisions, and health system or public health recommendations — a related template used by payers, health technology assessment bodies, and public health authorities, which adds criteria more relevant at a population or system level (budget impact at scale, and a more explicit priority-setting lens) while keeping the same certainty-of-evidence backbone.
A separate GRADE extension exists for evidence supporting a diagnostic test recommendation, since test accuracy studies and their downstream clinical consequences need a different evidence base than an intervention’s direct effect on outcomes. Anyone building an EtD table should confirm which template matches their decision type before starting — using the clinical-recommendation criteria for what is actually a coverage decision will leave out considerations reviewers and adopting bodies expect to see.
Common pitfalls when building or reading an EtD table
- Treating certainty of evidence as the whole story. A low-certainty rating does not automatically produce a conditional recommendation, and a high-certainty rating does not automatically produce a strong one — a large, unambiguous balance of effects can still support a strong recommendation on moderate-certainty evidence, and a values criterion with wide variability can pull a high-certainty finding down to conditional.
- Skipping the “additional considerations” free-text field. The judgment columns compress a lot of nuance into a short label; the accompanying research-evidence and additional-considerations notes are where a panel’s actual reasoning lives, and they are what a reader needs to evaluate whether the panel’s judgment was sound.
- Confusing an EtD framework with the certainty-of-evidence rating alone. A guideline that reports “GRADE: moderate” without a documented EtD table has not shown its work on how that certainty rating became a specific recommendation — the two are related but distinct outputs, and only the second one produces an actionable recommendation.
Illustrative composite, not a real guideline panel: imagine a panel assessing a moderate-certainty finding that a home-monitoring intervention modestly reduces hospital readmissions. Desirable effects are real but modest; undesirable effects (alert fatigue, false positives) are judged small; resource requirements are substantial for smaller health systems; equity considerations favor the intervention where broadband access is reliable and disfavor it where it isn’t. That combination — real but modest benefit, uncertain resourcing, uneven feasibility — is a textbook case for a conditional recommendation with an explicit note on the access dependency, rather than either a strong recommendation or no recommendation at all.
Frequently asked questions
Is an EtD framework the same thing as a GRADE certainty rating?
No. The certainty rating (high/moderate/low/very low) is one input into the EtD framework, not the framework itself. The EtD table adds values, balance of effects, resources, equity, acceptability, and feasibility on top of the certainty rating to reach a recommendation.
Do I need GRADEpro GDT to build an EtD framework?
No specific software is required by the GRADE methodology itself, but GRADEpro GDT is the tool most guideline-development groups use because it structures the criteria, judgment options, and evidence tables in the standard published format, which makes the output easier for other panels and journals to recognize and review.
What’s the difference between a strong and a conditional recommendation?
A strong recommendation means the panel is confident the balance of consequences clearly favors one option — most people would want it, and it can reasonably anchor a default or a quality measure. A conditional (weak) recommendation means the trade-offs are closer or values vary, and it should prompt individualized, shared decision-making rather than a blanket default.
Does the same EtD template apply to a coverage decision as to a clinical guideline?
No. GRADE publishes separate, related templates: one for clinical recommendations aimed at clinicians and patients, and another for coverage decisions and health-system/public-health recommendations, which weight population-level and budget-impact considerations differently. Confirm which template fits the decision before building the table.








