Skip to main content
v2026.11,772 entries · CC-BY 4.0

PHQ-9: The Nine Items, Scoring Bands, and Why Item 9 Is Not a Risk Assessment

The PHQ-9 scores nine DSM-mapped depression items from 0 to 27. This guide covers the item set, both scoring methods, the severity bands, what the diagnostic accuracy evidence actually supports, and the escalation obligations item 9 creates.

Written and maintained by CASRAI Editorial Board

Last updated

The Patient Health Questionnaire-9 (PHQ-9) is a nine-item self-report instrument that measures the severity of depressive symptoms over the preceding two weeks. It was developed by Kroenke, Spitzer and Williams as the depression module of the PRIME-MD diagnostic instrument, and published as a standalone measure in the Journal of General Internal Medicine in 2001. It is now the most widely deployed depression measure in general medical settings worldwide, embedded in electronic health records, primary-care workflows, integrated behavioural-health programmes and clinical-trial outcome batteries.

Its design is unusual among rating scales in one specific respect: the nine items are not an empirically derived factor structure, they are a near-verbatim restatement of the nine DSM criterion-A symptoms for a major depressive episode. That one-to-one mapping is what allows the same nine responses to be scored two entirely different ways — as a continuous severity score, or as a provisional diagnostic algorithm. Most implementations use only the first and never configure the second.

This guide covers the item set and its DSM mapping, both scoring methods, the severity bands, what the diagnostic accuracy literature actually supports (which is meaningfully less than the figure most training materials quote), and the item that generates the majority of the operational risk for hospital patient-safety and quality teams: item 9.

The nine items and their DSM-5 mapping

The stem for all nine items is: “Over the last 2 weeks, how often have you been bothered by any of the following problems?” Each item corresponds to one DSM-5 criterion-A symptom of a major depressive episode.

# Item (abbreviated) DSM-5 criterion-A symptom
1 Little interest or pleasure in doing things Anhedonia
2 Feeling down, depressed, or hopeless Depressed mood
3 Trouble falling or staying asleep, or sleeping too much Sleep disturbance (insomnia or hypersomnia)
4 Feeling tired or having little energy Fatigue or loss of energy
5 Poor appetite or overeating Appetite or weight change
6 Feeling bad about yourself, or that you are a failure, or have let yourself or your family down Worthlessness or excessive guilt
7 Trouble concentrating on things, such as reading the newspaper or watching television Diminished concentration or indecisiveness
8 Moving or speaking so slowly that other people could have noticed — or the opposite, being so fidgety or restless that you have been moving around a lot more than usual Psychomotor retardation or agitation
9 Thoughts that you would be better off dead, or of hurting yourself in some way Recurrent thoughts of death or suicidal ideation

Items 1 and 2 correspond to the two DSM “gateway” symptoms — at least one of which must be present for a major depressive episode. Items 3 through 8 are drawn directly from the criterion list. Item 8 is the only double-barrelled item, covering two opposite psychomotor presentations in a single response; a patient who is retarded on some days and agitated on others has no clean way to answer it.

The tenth question, which is not scored

The standard PHQ-9 form carries a tenth question after the nine items: “If you checked off any problems, how difficult have these problems made it for you to do your work, take care of things at home, or get along with other people?” with four options from “not difficult at all” to “extremely difficult.” This is a functional-impairment item, and it is not added to the 0–27 total. It exists because DSM requires clinically significant distress or impairment (criterion B) in addition to the symptom count. Implementations that drop it lose the only impairment signal on the form; implementations that accidentally sum it produce scores that will not compare against any published band or cutoff.

How the PHQ-9 is scored

Each of the nine items is scored 0 to 3 against a frequency-anchored response set:

  • 0 — Not at all
  • 1 — Several days
  • 2 — More than half the days
  • 3 — Nearly every day

The total is the simple unweighted sum, giving a range of 0 to 27. There is no reverse scoring, no subscale weighting and no age or sex adjustment. This arithmetic simplicity is a large part of why the instrument spread as fast as it did — it can be scored by hand at the bedside in seconds, and by an EHR flowsheet with no logic beyond addition.

Because the anchors are frequency-based rather than intensity-based, the score is a statement about how often symptoms occurred, not how severe they felt. A patient with profoundly intense but intermittent symptoms can score lower than one with mild, constant ones. This is a design choice inherited from the DSM criteria, which are themselves largely frequency- and duration-defined, but it is worth understanding before treating the total as a pure intensity measure.

Handling missing items

The conventional rule is that a PHQ-9 with more than two missing items should not be scored. With one or two missing, the common approach is prorating — take the mean of the completed items and multiply by nine. Whatever rule a site adopts, it needs to be written down and applied consistently, because a prorated score and a partial-sum score for the same patient can differ by several points and land in different severity bands. Silent treatment of blanks as zeroes is the most common and most damaging error: it systematically biases scores downward, and it does so most often on item 9, which is the item patients skip.

Severity bands

The conventional interpretation bands, unchanged since the original publication, are:

Total score Depression severity Conventional action
0–4 Minimal or none No treatment indicated on the basis of the score
5–9 Mild Watchful waiting; repeat at follow-up
10–14 Moderate Treatment plan — counselling, follow-up, pharmacotherapy considered
15–19 Moderately severe Active treatment with pharmacotherapy and/or psychotherapy
20–27 Severe Immediate initiation of pharmacotherapy; expedited specialist referral

Two things about these bands are routinely misread. First, they are severity descriptors, not diagnoses — “moderately severe depression” on a PHQ-9 is a description of a symptom score, not a clinical diagnosis of major depressive disorder. Second, the “conventional action” column is a proposal from the instrument’s authors, not a regulatory requirement or a standard of care. Sites that hard-wire it into an EHR as a forcing rule should understand they are adopting an editorial recommendation, not implementing a guideline.

The ≥10 cutoff is the one that carries real evidentiary weight, because it is the threshold the diagnostic-accuracy literature has actually tested at scale. The boundaries at 5, 15 and 20 are conventional divisions of the range and have far less validation behind them.

The two scoring methods: cutoff versus algorithmic

The same nine responses support two distinct scoring approaches, and they do not identify the same patients.

Cutoff (severity) scoring sums the items and compares the total against a threshold, normally ≥10. This is what virtually every EHR implementation does.

Algorithmic (diagnostic) scoring ignores the total and applies the DSM criteria structure directly: count how many items are endorsed at “more than half the days” or greater (item 9 counts if endorsed at all, at any frequency), and require that at least one of the endorsed items is item 1 or item 2. Five or more such symptoms suggests probable major depressive disorder; two to four suggests probable other depressive disorder.

The two methods disagree in both directions. A patient endorsing all nine items at “several days” scores 9 on cutoff scoring — below threshold — but has no items at criterion frequency and so is also negative algorithmically. A patient endorsing four items at “nearly every day” scores 12 and crosses the cutoff, but has only four criterion-level symptoms and does not meet the algorithm. In practice the algorithmic method is more specific and less sensitive. Sites should know which one their build uses, because a quality measure defined against one and reported from the other will not reconcile.

How accurate the PHQ-9 actually is

This is where most institutional training material is out of date, and it matters for anyone setting screening policy or writing a measure specification.

The original 2001 validation study reported sensitivity of 88% and specificity of 88% for major depression at a cutoff of ≥10. That paired figure is the one still reproduced on most reference cards, intranet pages and vendor collateral.

The individual participant data (IPD) meta-analysis by Levis and colleagues, published in the BMJ in April 2019 (BMJ 2019;365:l1476), pooled primary data across a large set of studies and produced a materially lower estimate. Among studies using a semi-structured diagnostic interview as the reference standard, pooled sensitivity at ≥10 was 0.85 (95% CI 0.79 to 0.89) and pooled specificity 0.85. The analysis confirmed that ≥10 maximises combined sensitivity and specificity overall and across subgroups — so the conventional cutoff survives — but the accuracy at that cutoff is lower than the original figure, and IPD meta-analysis is a stronger design than the single-sample validation it supersedes. An updated systematic review and IPD meta-analysis from the same research programme was published in 2021.

The practical consequence is a base-rate problem, not a subtle one. Specificity of 0.85 means roughly 15% of people without major depression screen positive. In a general medical population where the prevalence of major depression may be under 10%, the majority of positive PHQ-9 screens will not be major depression on structured interview. That is not an argument against screening — it is an argument that a positive PHQ-9 obligates a diagnostic follow-up step, and that a screening programme without the capacity to deliver that follow-up generates work it cannot discharge. This is exactly the “adequate systems in place” qualifier that accompanies mainstream screening recommendations.

Reporting the two figures side by side is worth doing in local training material. Staff who have been taught 88/88 and then encounter a false-positive rate consistent with 85 tend to conclude the tool is broken, rather than that their prior was wrong.

Item 9, suicidality, and the escalation problem

Item 9 — “Thoughts that you would be better off dead, or of hurting yourself in some way” — is the single highest-risk element of the instrument from a patient-safety and risk-management standpoint, and the one most likely to be implemented badly.

Item 9 is not a suicide risk assessment

This is the central point and it cannot be softened. Item 9 is one frequency-rated question about passive death wish or self-harm ideation, conflated into a single response. It does not ask about intent, plan, means, preparatory behaviour, prior attempts, or protective factors — the elements that actually differentiate risk levels. A patient scoring 1 on item 9 and a patient with a formulated plan and access to means can produce the same value on the form.

Treating a non-zero item 9 as a completed risk assessment is a documented failure mode. Its inverse — treating a zero on item 9 as evidence that a patient is not at risk — is worse, because it converts a screening instrument into false reassurance that gets charted. Item 9 is a trigger for assessment. It is never the assessment.

Where a structured suicide risk assessment is required, hospitals use purpose-built instruments — the Columbia-Suicide Severity Rating Scale (C-SSRS) and the Ask Suicide-Screening Questions (ASQ) toolkit are the two most commonly adopted — followed by a clinician risk formulation. The PHQ-9 feeds that pathway; it does not replace any part of it.

The closed-loop requirement

The operational risk in PHQ-9 deployment is almost never the questionnaire. It is the gap between a positive item 9 and a human being responding to it. Common and recurring failure patterns:

  • Asynchronous administration with no monitoring. PHQ-9s pushed to a patient portal, kiosk or tablet before an appointment can surface a positive item 9 into a queue nobody watches in real time — or outside clinic hours entirely. Any asynchronous channel needs a defined monitoring window and a stated response time, or it should not carry item 9.
  • Scored total without item-level visibility. A build that stores only the 0–27 total makes a positive item 9 invisible to anyone reviewing the chart. Item 9 must be independently retrievable and independently alertable, not buried in a sum. A patient can score 6 — “mild” — with a positive item 9.
  • Alert with no owner. A best-practice advisory that fires to whoever happens to have the chart open is not an escalation pathway. The policy needs a named role, a defined timeframe, and a documented action.
  • No environmental follow-through. A patient identified as at risk may need environmental controls alongside clinical assessment — see ligature risk assessment for the physical-environment side and elopement risk assessment for the egress side.
  • Screening in settings without the follow-up capacity. Deploying item 9 in a service with no behavioural-health pathway creates a documented positive finding with no documented response. That record is discoverable.

Suicide-risk reduction is an explicit element of the Joint Commission’s National Patient Safety Goals, which require validated screening and a documented response pathway for the relevant patient populations. A PHQ-9 build that captures item 9 but cannot demonstrate what happened next will not satisfy that expectation, and a suicide of a patient in a staffed care setting falls within the scope of events tracked as serious reportable events. The screening instrument is the easy part of compliance; the closed loop is the part surveys actually examine.

The PHQ-8 alternative

The PHQ-8 is the PHQ-9 with item 9 removed, scored 0–24. It was developed for population and epidemiological research where administering a suicidality item by post or telephone, with no clinician able to respond, is itself an ethical problem. A systematic review and IPD meta-analysis published in Psychological Medicine found the diagnostic accuracy of the PHQ-8 and PHQ-9 to be equivalent — item 9 contributes very little discriminative information for detecting major depression, because it is endorsed relatively rarely.

That finding is genuinely useful for programme design. If a survey, registry or research protocol cannot guarantee a real-time clinical response, the defensible choice is the PHQ-8, not the PHQ-9 with a disclaimer. Removing the item costs essentially nothing in screening performance and removes the obligation entirely. Conversely, in any clinical setting where a response pathway exists, item 9 should be retained — its value there is as a safety trigger, not as a contributor to the depression score.

PHQ-2, PHQ-4 and the rest of the family

  • PHQ-2 — items 1 and 2 only (anhedonia and depressed mood), scored 0–6. Used as a first-stage screen with a conventional cutoff of ≥3, followed by the full PHQ-9 if positive. This two-stage design substantially reduces administration burden in high-volume settings while preserving sensitivity.
  • PHQ-8 — the nine-item form without item 9, scored 0–24, for research and population surveillance.
  • GAD-7 — the seven-item Generalized Anxiety Disorder scale from the same research group, scored 0–21, and the near-universal companion measure to the PHQ-9. Co-administration is standard in integrated behavioural health, because depression and anxiety are highly comorbid and the two instruments together take under five minutes.
  • PHQ-4 — an ultra-brief four-item combined screen: PHQ-2 items plus the first two GAD-7 items, scored 0–12.
  • PHQ-15 — a somatic symptom severity measure from the same instrument family, distinct in construct from the depression module.
  • PHQ-9 Modified for Adolescents (PHQ-A) — an adolescent adaptation with modified wording and additional suicidality items.

The PHQ-9 as an outcome measure

Beyond screening, the PHQ-9 is one of the most frequently used outcome measures in depression clinical trials and in measurement-based care programmes. Three properties drive that:

It is sensitive to change. The 0–27 range with frequency anchors gives enough resolution to track response over repeated administrations, which is the core mechanic of measurement-based care and of collaborative-care models generally.

It has conventional response definitions. Trials and registries commonly define response as a ≥50% reduction from baseline and remission as a score below 5. A change of around 5 points is widely treated as the threshold for a clinically meaningful difference, though the appropriate value depends on the population and on the estimation method — see anchor-based versus distribution-based MCID estimation for why a single universal number should be treated with caution.

It is free and ubiquitous. No licence fee, no vendor dependency, and translations into a large number of languages already exist — though a translation being available is not the same as it being validated in the target population, which is a distinction worth checking; see translation and back-translation of research instruments.

As a patient-reported instrument it sits inside the broader PROM landscape in hospital quality reporting, and the selection criteria in choosing and validating a PROM apply to it as they do to any other patient-reported measure.

Limitations

  • It is a screening and severity instrument, not a diagnostic one. A PHQ-9 score of 22 is not a diagnosis of severe major depressive disorder. Diagnosis requires clinical interview, differential consideration (bipolar disorder, substance-induced mood disorder, medical causes such as hypothyroidism or anaemia, bereavement, adjustment disorder) and assessment of impairment and duration. The instrument cannot distinguish unipolar from bipolar depression at all — a critical gap, because antidepressant monotherapy in undiagnosed bipolar disorder carries real harm.
  • Somatic item confounding in medically ill populations. Items 3, 4, 5 and 7 — sleep, energy, appetite, concentration — are routinely elevated by physical illness, hospitalisation, chemotherapy, chronic pain and the inpatient environment itself. In oncology, post-stroke, dialysis, post-surgical and general inpatient populations the PHQ-9 will over-identify. This is the single largest source of false positives in hospital use, and it is a structural property of the instrument rather than an implementation error.
  • Self-report vulnerability. Scores are affected by health literacy, language, cognitive impairment, delirium and deliberate under-reporting where the patient perceives a consequence — loss of a licence, custody implications, employment or immigration concerns, or admission. In patients with cognitive impairment or acute confusion, self-report validity degrades sharply; see CAM-ICU and Mini-Cog for the assessment of the confounding conditions themselves.
  • Floor effects in mild populations. In samples with low symptom burden a large share of respondents cluster at or near zero, compressing variance and limiting the ability to detect improvement — see floor and ceiling effects.
  • Item 8 is double-barrelled. Psychomotor retardation and psychomotor agitation are opposite presentations sharing one response option.
  • Frequency, not intensity. The anchors count days, not severity, so the total does not straightforwardly represent how bad the symptoms are.
  • Repeated administration effects. Scores tend to decline on retest independent of treatment, through regression to the mean and familiarity with the instrument. This inflates apparent improvement in single-arm programme evaluations that lack a comparison group. On the general question of what a repeat administration should and should not be expected to reproduce, see test-retest reliability.

Implementation notes for quality and patient-safety teams

  • Store item-level responses, not just the total. This is the prerequisite for item 9 alerting, for algorithmic rescoring, and for any subsequent measure or research use. Totals-only builds are extremely difficult to remediate after the fact.
  • Define the missing-data rule in policy and make sure the EHR implements the same rule the analytics layer does. Blanks silently coerced to zero is the default failure.
  • Route item 9 independently of the total. The alert condition is a non-zero item 9, not a total above a threshold.
  • Specify which scoring method is in use in any measure definition, and do not let the screening build and the reporting build diverge.
  • Match the form to the response capacity. PHQ-8 where there is no real-time clinical response; PHQ-9 where there is.
  • Record administration mode and language. Self-administered, interviewer-administered, telephone and portal administrations are not freely interchangeable, and mode changes mid-programme confound trend analysis.
  • Audit the loop, not the screening rate. Screening completion rates are easy to hit and prove very little. The defensible metric is the proportion of positive item 9 responses with a documented assessment inside the policy timeframe.

Within a hospital’s wider assessment-instrument portfolio, the PHQ-9 sits alongside other validated, item-scored screens with defined cutoffs and defined escalation consequences — the Braden Scale for pressure-injury risk, the Morse Fall Scale for falls, the RASS for sedation depth, and the Clinical Frailty Scale for frailty. The same governance question applies to all of them: not whether the score is captured, but whether the action the score is supposed to trigger reliably happens. Broader programme context sits on the patient safety pillar.

Copyright and permitted use

The PHQ-9 is free to use. Pfizer, which funded the original PRIME-MD development work, released the PHQ family of instruments and the GAD-7 into the public domain in 2010, with no copyright restriction and no charge. No permission, licence or fee is required to reproduce, translate, display or distribute it, including within commercial software.

Many copies still in circulation carry a “Copyright © Pfizer Inc.” line and a PRIME-MD trademark notice that predate that release. Those notices are historical and do not restrict use. This matters practically: sites occasionally route PHQ-9 EHR builds or patient-facing materials through a licensing review that is not required, or pay a vendor for access to an instrument that is free.

Frequently asked questions

What is a normal PHQ-9 score?

A total of 0–4 is the minimal or no-depression band. Scores of 5–9 indicate mild symptoms, 10–14 moderate, 15–19 moderately severe and 20–27 severe. A score is a severity descriptor, not a diagnosis.

What PHQ-9 score indicates depression?

≥10 is the conventional screening cutoff, and the pooled IPD meta-analytic evidence confirms it maximises combined sensitivity and specificity. A score at or above 10 indicates a positive screen requiring diagnostic follow-up, not a diagnosis of depression in itself.

How accurate is the PHQ-9?

At ≥10, the 2019 IPD meta-analysis by Levis and colleagues found pooled sensitivity of 0.85 (95% CI 0.79 to 0.89) and pooled specificity of 0.85 against semi-structured interview reference standards. The original 2001 validation study reported 88% for both, which is higher than the pooled estimate and is still widely quoted.

What does a positive item 9 require?

A documented suicide risk assessment by a qualified clinician within a timeframe defined in local policy. Item 9 is a trigger, not an assessment — it does not capture intent, plan, means or prior attempts, and a zero on item 9 is not evidence of absent risk.

What is the difference between the PHQ-9 and the PHQ-8?

The PHQ-8 omits item 9 (suicidality) and is scored 0–24. Their diagnostic accuracy for major depression has been found equivalent in IPD meta-analysis, so the PHQ-8 is the appropriate choice for surveys and research where no clinician can respond to a positive suicidality item in real time.

What is the difference between the PHQ-9 and the GAD-7?

The PHQ-9 measures depressive symptom severity across nine items (0–27); the GAD-7 measures generalised anxiety symptom severity across seven items (0–21). They come from the same research group, share the same four-point frequency anchors and two-week recall window, and are routinely administered together.

Is the PHQ-9 diagnostic?

No. It is a screening and severity-monitoring instrument. Its algorithmic scoring method yields a provisional diagnostic impression aligned to DSM criteria, but diagnosis requires clinical interview, differential consideration and impairment assessment. Notably, it cannot distinguish unipolar from bipolar depression.

Does the PHQ-9 require a licence or fee?

No. It was placed in the public domain in 2010 and requires no permission or payment to use, translate or distribute, including in commercial products. Copyright notices on older printed copies predate that release.

How often should the PHQ-9 be repeated?

There is no universal interval. In measurement-based care it is commonly repeated at each treatment contact or monthly during active treatment to track response, defined conventionally as a ≥50% reduction from baseline, and remission as a score below 5. Sites should be aware that scores decline on retest independent of treatment effect.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about PHQ-9: The Nine Items, Scoring Bands, and Why Item 9 Is Not a Risk Assessment

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.