Skip to main content
v2026.11,772 entries · CC-BY 4.0

Choosing and Validating a Patient-Reported Outcome Measure (PROM)

A five-step checklist for choosing an existing patient-reported outcome measure rather than building one: matching the construct, finding candidates, checking population-specific validation evidence, confirming an MCID exists, and clearing licensing and translation.

Written and maintained by CASRAI Editorial Board

Last updated

On this page: a five-step checklist for choosing an existing patient-reported outcome measure (PROM) rather than building one — matching the construct and context of use, finding candidate instruments, checking that the validation evidence actually applies to your population, confirming an MCID exists for the change you want to detect, and clearing licensing and translation before you commit to an instrument.

This Guide Is About Selecting an Instrument, Not Building or Qualifying One

Most studies that need a patient-reported outcome do not need a new instrument — they need the right existing one. Building a new PROM from scratch means running the full measurement-development sequence (concept elicitation, item generation, cognitive interviewing, field testing, psychometric evaluation) before you can collect a single usable data point, and it is rarely justified when a validated instrument already covers the construct. This guide is a practical selection checklist for a researcher choosing among existing instruments: what to check before you commit, in what order, and where it can go wrong.

Two adjacent questions this guide does not answer:

  • If you are building the regulatory evidence package to support a labeling claim — context of use, FDA’s Clinical Outcome Assessment (COA) Qualification Program, the difference between an instrument being validated and being qualified — see Clinical Outcome Assessment Validation.
  • If you are selecting or reporting a PROM specifically for CMS hospital quality measures (the THA/TKA PRO-PM collection windows and response-rate thresholds), see PROMs in Hospital Quality Reporting.

Both of those are heavier processes than most study teams need. This checklist is for the more common case: a study, dissertation, or program evaluation that needs to measure a construct in a defined population and wants to use a real, defensible, already-validated instrument to do it.

Step 1: Define the Construct and the Context of Use Before You Search

The single most common selection mistake is searching for an instrument before precisely defining what needs to be measured. “Quality of life,” “pain,” and “functional status” are each umbrella terms covering dozens of non-interchangeable instruments built around different underlying constructs (pain intensity vs. pain interference; generic health-related quality of life vs. disease-specific quality of life). Write down, before you search:

  • The concept of interest — the specific construct, stated narrowly (e.g., disease-specific physical function, not generic health status).
  • The target population — condition, severity range, age band, language(s), literacy level, and any subgroup the instrument must work equally well across.
  • Timing and mode of administration — single time point or repeated measures, self-administered on paper, electronic (see ePRO), interview-administered, or proxy-reported.
  • What the score needs to support — a between-group comparison, an individual-level change score, or a screening cutoff. These place different demands on the instrument’s measurement properties, particularly responsiveness and measurement error.

This is a lighter-weight version of the same idea FDA formalizes as context of use for regulatory COAs — the same instrument can be well-supported for one context of use and unsupported for another. Writing your context of use down first turns instrument selection from a keyword search into a matching exercise against a fixed specification.

Step 2: Build a Candidate List From Instrument Registries, Not a General Web Search

A general search tends to surface whichever instrument is most cited, not the one that best fits your construct and population. Better starting points:

  • PROMIS and the other NIH-funded item banks, hosted through HealthMeasures — item banks covering physical, mental, and social health domains, available as fixed-length short forms or computerized adaptive tests (see Item Response Theory for how CAT scoring works). Broad population coverage and, notably, no licensing fee.
  • The COMET Initiative database (comet-initiative.org) — core outcome sets agreed by consensus for specific conditions and trial types, useful for checking what outcome domains and instruments a field has already converged on.
  • The COSMIN database of systematic reviews (cosmin.nl) — systematic reviews of measurement properties, organized by construct and population, that summarize which instruments have (and have not) been evaluated in your population already, sparing you a from-scratch literature search.
  • Condition-specific instrument developers and registries — e.g., the EORTC Quality of Life Group for oncology, EuroQol for the EQ-5D family — when the construct is disease-specific rather than generic.

Aim for three to five candidates before moving to Step 3, not one. A shortlist lets you compare validation evidence and licensing terms directly instead of anchoring on the first instrument you find.

Step 3: Check That the Validation Evidence Applies to Your Population — Not Just That It Exists

An instrument being “validated” is not a fixed, portable property. Validity evidence is specific to a population, language, and context — an instrument well-validated in English-speaking adults with moderate osteoarthritis carries no automatic evidence that it performs the same way in adolescents, in a different language, or at a different disease severity. The COSMIN taxonomy (Mokkink et al., Journal of Clinical Epidemiology, 2010) is the standard framework for checking this systematically, organizing measurement properties into three domains:

  • Reliability — internal consistency, test–retest reliability (see Test-Retest Reliability), and measurement error (see Minimal Detectable Change).
  • Validity — content validity (does it cover the right content, per your population’s own input), structural validity, cross-cultural validity/measurement invariance (does the instrument mean the same thing across the language or cultural groups in your study), and criterion or construct validity.
  • Responsiveness — whether the instrument detects change over time in the direction and magnitude a true clinical change would produce, which matters specifically if your design measures change rather than a single time point.

For each candidate, ask directly: has this instrument been evaluated in a population like mine on the properties my design depends on? An instrument with excellent internal consistency but no evidence of responsiveness is a poor choice for a pre/post design even if it is well-validated overall. See Psychometrics for how these properties are estimated, and Cronbach’s Alpha in SPSS for the internal-consistency calculation specifically.

Step 4: Confirm an MCID Exists for Your Population and Change Metric

If the study’s purpose includes judging whether a change is clinically meaningful — not just statistically detectable — the instrument needs a published minimal clinically important difference (MCID) or minimal important change (MIC) value for a population reasonably close to yours. This is a separate check from Step 3’s general validation evidence, and it is easy to skip because an instrument can be well-validated with no MCID published for it at all, or with an MCID established only in a population that doesn’t match your study.

Two things worth confirming before you commit:

  • Whether the published MCID was derived by anchor-based methods (typically a patient global rating of change), distribution-based methods, or both — see MCID: Anchor-Based vs. Distribution-Based Estimation for how each is calculated and what each assumes.
  • Whether the MCID is being confused with the instrument’s minimal detectable change (MDC) — a statistical noise threshold, not a clinical-importance threshold. The two answer different questions and an observed change needs to clear both to be interpreted as a real, meaningful improvement; see Minimal Detectable Change for the distinction and the decision rule that combines them.

An MCID established in one population, language, or disease-severity band does not automatically transfer to another — the same caveat as Step 3’s validation evidence. If no MCID exists for anything close to your population, that is a real limitation to plan around (a larger sample sized on statistical detectability alone, or a companion anchor question built into your own protocol) rather than one to discover after data collection.

Step 5: Clear Licensing, Cost, and Translation Before You Commit

Licensing terms vary enormously across PROMs and are easy to discover too late — after a protocol is already written around a specific instrument. Broadly, instruments fall into two categories:

  • Freely available, no license fee. PROMIS and the other NIH-funded HealthMeasures item banks are available for research and clinical use at no cost, which is one reason they are a reasonable default starting point in Step 2.
  • Copyright-holder licensed. Many widely used instruments require a license agreement with the developer or a licensing agent before use, and terms differ for academic, non-commercial, and industry-sponsored studies — for example, the SF-36/SF-12 family (licensed through Optum), the EQ-5D family (licensed through EuroQol), and the EORTC QLQ instruments (registered and, depending on the study type, licensed through the EORTC Quality of Life Group). Fees, registration requirements, and permitted uses (paper vs. electronic administration, modification, redistribution) vary by instrument and by license type, and change over time — confirm current terms directly with the rights holder rather than from a secondary source, and budget the time to do this before finalizing a protocol, not after.

If the study population includes non-English speakers, also confirm a validated translation already exists in the needed language(s) before assuming one does. A validated instrument in English is not automatically validated once translated — a good translation still needs to go through forward translation, back-translation, reconciliation, and cognitive debriefing (the process ISPOR’s Good Practice principles for PRO translation and cultural adaptation set out) before the translated version carries the same validity evidence as the original. Licensing terms typically govern whether you may commission your own translation at all, or whether a validated translation must be licensed from the rights holder directly — check both.

Quick-Reference Selection Checklist

  • Concept of interest, target population, timing/mode, and what the score needs to support are written down before searching.
  • Three to five candidate instruments identified from a registry (HealthMeasures/PROMIS, COMET, COSMIN database, or a condition-specific developer), not a single general search result.
  • For each candidate: validation evidence checked against COSMIN’s reliability, validity, and responsiveness domains — specifically in a population comparable to yours, not just “a” population.
  • A published MCID or MIC exists for a population close to yours, and you know whether it was derived by anchor-based or distribution-based methods.
  • Licensing terms confirmed directly with the rights holder (or confirmed free-to-use, e.g. PROMIS) for your specific use case — academic, clinical, or industry-sponsored.
  • If the population is multilingual: a validated translation exists in the needed language(s), or a properly licensed translation process is budgeted into the timeline.

Frequently Asked Questions

Is validation in a general adult population enough, or do I need population-specific evidence?

General-population validation is a starting point, not a substitute for population-specific evidence if your study population differs meaningfully — a different age band, disease severity, language, or care setting. Check Step 3’s COSMIN domains specifically against your population before assuming general validity transfers.

Can I modify or shorten a PROM’s wording to fit my study?

Not without consequence to its validity evidence, and often not without violating the license. Any wording change, item drop, or response-scale change should be treated as creating a new, unvalidated instrument unless the developer has published and validated that specific modified version. If item burden is the concern, look for an official short-form version or, for item-bank-based instruments like PROMIS, a computerized adaptive test rather than editing items yourself.

What if no MCID exists for my exact population?

Use the closest population for which one has been established, report that as a limitation explicitly, and consider whether your protocol can build in an anchor (e.g., a patient global rating of change item) to derive a study-specific estimate. Don’t silently apply a general-population MCID to a substantially different population without flagging the assumption.

Is PROMIS really free to use?

PROMIS and the other NIH-funded HealthMeasures item banks are available for research and clinical use without a license fee, which is a genuine practical advantage over commercially licensed alternatives — confirm current registration requirements on HealthMeasures directly, since administrative requirements (account registration, attribution) can still apply even where no fee does.

Sources

  • U.S. Food and Drug Administration. Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims. Guidance for Industry, issued December 8, 2009.
  • Mokkink LB, Terwee CB, Patrick DL, Alonso J, Stratford PW, Knol DL, Bouter LM, de Vet HC. The COSMIN study reached international consensus on taxonomy, terminology, and definitions of measurement properties for health-related patient-reported outcomes. Journal of Clinical Epidemiology. 2010;63(7):737–745. DOI 10.1016/j.jclinepi.2010.02.006.
  • COSMIN taxonomy of measurement properties.
  • Jaeschke R, Singer J, Guyatt GH. Measurement of health status. Ascertaining the minimal clinically important difference. Controlled Clinical Trials. 1989;10(4):407–415. DOI 10.1016/0197-2456(89)90005-6.
  • de Vet HC, Terwee CB, Ostelo RW, Beckerman H, Knol DL, Bouter LM. Minimal changes in health status questionnaires: distinction between minimally detectable change and minimally important change. Health and Quality of Life Outcomes. 2006;4:54. DOI 10.1186/1477-7525-4-54.
  • Wild D, Grove A, Martin M, Eremenco S, McElroy S, Verjee-Lorenz A, Erikson P (ISPOR Task Force for Translation and Cultural Adaptation). Principles of Good Practice for the Translation and Cultural Adaptation Process for Patient-Reported Outcomes (PRO) Measures. Value in Health. 2005;8(2):94–104.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Choosing and Validating a Patient-Reported Outcome Measure (PROM)

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.