Skip to main content
v2026.11,772 entries · CC-BY 4.0

SF-36 (36-Item Short Form Health Survey) and Health-Related Quality of Life

How the SF-36 measures health-related quality of life across eight domains, how PCS/MCS summary scores are derived, and how SF-36 data is mapped to SF-6D or compared with EQ-5D to produce utility values for QALY-based health-economic models.

Written and maintained by CASRAI Editorial Board

Last updated

The SF-36 (36-Item Short Form Health Survey) is the most widely used generic instrument for collecting health-related quality of life (HRQoL) data in clinical research. It was developed as part of the RAND Medical Outcomes Study and validated in a landmark 1992 Medical Care paper by John E. Ware and Cathy Donald Sherbourne, which established the eight-scale structure the instrument still uses today. For a trialist or health economist, the practical question is rarely “what does the SF-36 measure” so much as “what can I actually do with the scores once I have them” — particularly when the endpoint that matters to a reimbursement body is a quality-adjusted life year (QALY), not a raw questionnaire score. This guide covers both: the instrument itself, and how its output is converted into the utility values that feed health-economic models.

What the SF-36 measures: eight domain scales

Ware and Sherbourne’s 1992 validation paper defined the SF-36 as measuring eight health concepts, each scored 0–100 with higher scores indicating better health:

  • Physical Functioning (PF) — limitations in physical activities such as walking, climbing stairs, or self-care, caused by health
  • Role-Physical (RP) — limitations in work or daily activities caused by physical health problems
  • Bodily Pain (BP) — the intensity of pain and its interference with normal work
  • General Health (GH) — personal evaluation of overall health, including resistance to illness
  • Vitality (VT) — energy versus fatigue
  • Social Functioning (SF) — the extent to which physical or emotional problems interfere with normal social activities
  • Role-Emotional (RE) — limitations in work or daily activities caused by emotional problems
  • Mental Health (MH) — psychological distress and well-being (anxiety, depression, positive affect)

These eight scales are not simply averaged into a single number. Each is scored independently, which is what makes the SF-36 useful as a profile instrument — it can show, for example, that an intervention improved physical functioning without materially affecting mental health, a distinction a single summary score would erase.

Physical and Mental Component Summary scores

For studies that need fewer than eight numbers to report or model, the eight scales can be aggregated into two higher-order scores: the Physical Component Summary (PCS) and the Mental Component Summary (MCS). These are derived through factor analysis of the eight domain scales — Physical Functioning, Role-Physical, Bodily Pain and General Health load most heavily onto the physical factor, while Mental Health, Role-Emotional, Social Functioning and Vitality load most heavily onto the mental factor (Bodily Pain, General Health, Social Functioning and Vitality carry meaningful cross-loadings on both, which is a known limitation of the two-summary reduction).

Both summary scores use norm-based scoring: they are transformed to a mean of 50 and a standard deviation of 10 against a general-population reference sample, so a score of 40 is one population standard deviation below average and a score of 60 is one above. This makes PCS/MCS interpretable across studies and populations in a way raw 0–100 domain scores are not, but it also means PCS/MCS are relative, norm-referenced numbers, not absolute quantities — they answer “how does this population compare to a reference population,” not “how much health does this person have” in any cardinal sense.

Why SF-36 scores are not utilities

This is the point where SF-36 data most often gets used incorrectly in economic submissions. The SF-36 (and its two component summaries) is a psychometric, profile-based instrument: its scores are ordinal-to-interval measures anchored to population norms, not a cardinal utility scale anchored at 0 (dead) and 1 (full health). QALY calculations require the latter — a utility value on the dead-to-full-health scale that can be multiplied by time to produce quality-adjusted life years. A PCS score of 45 cannot be plugged directly into a QALY formula; it has no defined relationship to the 0–1 utility scale a cost-utility model requires. Confusing “high SF-36 score” with “high utility value” is a common and avoidable error in study designs that were not planned with health-economic modeling in mind from the outset — a case for choosing the right patient-reported outcome measure before data collection begins, not after.

Getting from SF-36 to a utility: the mapping algorithm

Because so many trials collect SF-36 (or its shorter sibling, the SF-12) without a dedicated preference-based measure alongside it, a substantial literature exists on mapping SF-36 responses onto a preference-based index. The best-established route is the SF-6D, developed by Brazier, Roberts and Deverill (2002): a health-state classification derived from a subset of SF-36 items, reduced to six dimensions (physical functioning, role limitations, social functioning, pain, mental health and vitality) each with multiple levels, describing 18,000 distinct health states in total. A representative sample of the general population valued a subset of those states using the standard gamble technique, and the resulting model predicts a utility value for every one of the 18,000 possible SF-6D health states, anchored on the conventional 0–1 (dead-to-full-health) scale.

In practice this means: score the SF-36 responses into the SF-6D classification, look up (or model-predict) the corresponding utility value, and use that value in the cost-utility analysis. This is a legitimate, widely accepted route for producing utility estimates from a trial that only collected the SF-36 — but it is a modeled approximation, not a direct elicitation, and reviewers (NICE among them) expect it disclosed as such. A trial that anticipates needing QALYs is better served collecting a preference-based measure directly, and reporting a mapped SF-6D value as a sensitivity analysis rather than the base case.

SF-36/SF-6D versus EQ-5D for utility elicitation

The EQ-5D is the alternative most directly built for this purpose: it is a preference-based measure from the start, not a profile instrument later mapped onto one. Respondents self-classify their health state across five dimensions (mobility, self-care, usual activities, pain/discomfort, anxiety/depression, each with three or five severity levels depending on the version — EQ-5D-3L or EQ-5D-5L), and country-specific value sets, derived from general-population valuation studies, convert that classification directly into a utility index anchored on the 0–1 scale. There is no intermediate mapping step.

Consideration SF-36 → SF-6D mapping EQ-5D (direct)
Instrument type Profile measure, mapped after the fact Preference-based measure by design
Respondent burden 36 items 5 items (plus a VAS)
Utility precision Approximate; depends on mapping-algorithm fit Directly elicited, no approximation step
Descriptive richness Rich 8-domain clinical profile Coarser 5-dimension classification
HTA reference-case status Generally not preferred as base case (e.g., UK NICE reference case specifies EQ-5D) Widely specified as the reference-case default by HTA bodies

This is why the field has increasingly converged on EQ-5D for the specific job of utility elicitation, even in studies that also collect the SF-36 for its clinical and descriptive value. The two are not competitors for the same purpose: the SF-36 remains the stronger choice when the study needs a detailed, domain-level picture of physical and mental functioning for clinical interpretation, while EQ-5D (or another preference-based measure collected directly) is the stronger choice when the downstream requirement is a defensible utility value for a health-economic model. Many protocols now collect both: SF-36 for clinical/descriptive reporting, EQ-5D as the base-case utility source, with SF-6D-mapped SF-36 data retained as a sensitivity or bridging analysis. Economic evaluations built on either route should still follow the CHEERS 2022 reporting checklist, which explicitly requires disclosure of which outcome measure produced the utility values and how they were derived.

Practical implications for trial design

Three design decisions follow directly from the distinction above:

  • Decide early whether the trial needs QALYs. If a cost-utility analysis is anticipated (for a NICE technology appraisal or an equivalent HTA submission elsewhere), build in a direct preference-based measure such as EQ-5D from the outset rather than relying on a post hoc SF-36 mapping.
  • Collect SF-36 at the same timepoints as any clinical or patient-reported endpoints it will be interpreted alongside, and follow the same reporting discipline used for other patient-reported outcomes — see CONSORT-PRO and SPIRIT-PRO for how PRO data should be specified and reported in trial protocols and publications.
  • Report the mapping algorithm and version used if SF-6D-mapped values appear anywhere in the analysis (published mapping functions have been updated more than once since 2002) — an unspecified or outdated mapping function is a common point of challenge in HTA critique.

Frequently asked questions

Is the SF-36 itself a utility measure?

No. The SF-36 produces eight domain scores and, optionally, two norm-based component summaries (PCS and MCS). None of these are anchored on the 0 (dead) to 1 (full health) utility scale that QALY calculations require. A utility value has to be obtained either by collecting a preference-based measure directly (such as EQ-5D) or by mapping SF-36 responses onto a preference-based classification such as the SF-6D.

What is the difference between the SF-36 and the SF-12?

The SF-12 is a 12-item subset of the SF-36 designed to reproduce the Physical and Mental Component Summary scores with less respondent burden, at the cost of no longer supporting reliable scoring of all eight individual domain scales. Both can be mapped to the SF-6D.

Why does NICE (and similar HTA bodies) prefer EQ-5D over SF-36-derived utilities?

Because EQ-5D is elicited directly from a preference-based valuation exercise, whereas an SF-6D value mapped from SF-36 data is a modeled estimate with its own error and assumptions. HTA bodies generally treat the mapped value as an acceptable substitute when no direct preference-based measure was collected, not as an equally strong first choice.

Can SF-36 domain scores be used in a QALY calculation without going through SF-6D?

Not directly, and this is a common error. The 0–100 domain scores and the norm-based PCS/MCS scores have no defined mapping to the 0–1 utility scale on their own; some form of established mapping algorithm (SF-6D being the most widely used) is required to produce a value usable in a cost-utility model.

Does a higher SF-36 score always mean better health-related quality of life?

Within each of the eight scales, yes — all are scored so that 100 represents the best possible health state measured by that scale and 0 the worst. But because the scales are not utilities, “higher SF-36” does not translate proportionally into “more QALYs”; the relationship has to go through a mapping algorithm to be quantified.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about SF-36 (36-Item Short Form Health Survey) and Health-Related Quality of Life

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.