Measurement, Reliability & Validity
A study is only as good as its measurements. This sub-cluster covers the psychometric concepts that determine whether an instrument actually measures what it claims to: reliability (consistency of measurement, including Cronbach's alpha for internal consistency and test-retest reliability for stability over time) and validity (whether an instrument measures the intended construct, including content, construct, and criterion validity). It also covers instrument development and adaptation — how existing scales are chosen, modified, or built from scratch, and what evidence is needed to justify that choice to reviewers and funders.
Guides
Cohen’s Kappa for Two Raters: Formula and Worked Example
Cohen’s kappa is the chance-corrected agreement statistic for exactly two raters classifying nominal categories. Formula, a worked 2×2 example, Landis & Koch interpretation bands, and the kappa paradox explained.
Recall Bias: Causes, Direction, and How to Reduce It
Recall bias is differential memory error between cases and controls in a retrospective study. This guide covers why it happens, which way it can push a result, and the design fixes that actually reduce it.
Demand Characteristics: How Participants Guess Your Hypothesis and Change Their Behavior
Demand characteristics are cues that let participants guess a study’s hypothesis and adjust their behavior to match it. This guide covers Orne’s origin of the term, Weber and Cook’s four subject roles, how to detect it with a post-experimental inquiry, and how to design against it.
Test Equating and Linking
Equipercentile equating, IRT linking via Stocking-Lord and Haebara, and the anchor-item design rules that make scores from different test forms comparable.
Anchoring Vignettes: Correcting for Differential Item Functioning in Self-Reports
How anchoring vignettes correct differential item functioning in self-reported survey scales, using King and Hopkins’s response-consistency method.
Test Information Functions and Measurement Precision
Item information functions sum into a test information function that shows measurement precision at every trait level, not just one sample-average reliability number. This guide covers the IRT formulas, the SE(theta) = 1/sqrt(I(theta)) relationship, how to read a TIF curve, and how to use one to target a short form or evaluate an instrument at the range that matters.
Factor Scores: Methods and When Not to Use Them
Regression, Bartlett, and Anderson-Rubin factor-score methods compared with a reproducible worked example of factor-score indeterminacy, plus the specific cases where a sum or mean score is the better choice.
Item-Total Correlation in Scale Purification
The corrected item-total correlation diagnostic for scale purification: what it measures, the conventional 0.30 cutoff, and how it relates to alpha if item deleted, with a reproducible worked example.
Differential Item Functioning (DIF): Detecting Item-Level Bias with the Mantel-Haenszel Method
Differential item functioning (DIF) is an item-level fairness problem distinct from overall test bias or group score differences: it occurs when people at the same trait level respond to one item differently depending on group membership. This guide covers uniform vs. non-uniform DIF and walks through the Mantel-Haenszel detection procedure with a reproducible simulated example.
Floor and Ceiling Effects: Detecting Range Compression and Fixing It
How floor and ceiling effects compress variance at a scale’s extremes, distort correlations, group comparisons and responsiveness, and the instrument-selection and statistical fixes for each.
The Rasch Model: Item-Invariant Measurement and the Person-Item Map
The Rasch model is a one-parameter IRT model built around invariant measurement: item difficulty and person ability that hold constant regardless of who or what was used to estimate them. This guide covers specific objectivity, infit/outfit fit statistics, and how to read a person-item (Wright) map.
Convergent and Discriminant Validity
Convergent validity means a measure correlates with other measures of the same construct; discriminant validity means it does not correlate with measures of unrelated constructs. This guide covers both, the Campbell-Fiske multitrait-multimethod matrix that tests them together, and the modern AVE/Fornell-Larcker/HTMT thresholds.
Choosing and Validating a Patient-Reported Outcome Measure (PROM)
A five-step checklist for choosing an existing patient-reported outcome measure rather than building one: matching the construct, finding candidates, checking population-specific validation evidence, confirming an MCID exists, and clearing licensing and translation.
Fleiss’ Kappa for Multiple Raters: Formula and Worked Example
The Fleiss’ kappa formula for three or more raters, a fully worked five-subject rating-matrix example, why it isn’t simply Cohen’s kappa for 3+ raters, and the Landis & Koch bands carried over from Cohen’s kappa.
Krippendorff’s Alpha: Calculating Intercoder Reliability
The observed-vs-expected disagreement formula, a worked example with missing data across three coders, Krippendorff’s own 0.667/0.800 benchmarks, a bootstrap confidence-interval method, and software (R, Python, SPSS/SAS/Stata, ReCal).
Item Response Theory for Research Scales
What item response theory is, how the item-characteristic curve and its difficulty and discrimination parameters work, what the 1PL, 2PL, and 3PL models add, and why IRT (unlike classical test theory) enables adaptive testing.
Split-Half Reliability and the Spearman-Brown Correction
A split-half correlation is the reliability of a half-length test, not the full one, and different equally valid splits of the same items produce different numbers. What the Spearman-Brown correction fixes, what it can’t fix, and why Cronbach’s alpha — the average of every possible split — became the standard instead.
Bland-Altman Plots for Method Agreement: Why Correlation Isn’t the Right Tool
Why Pearson correlation cannot assess agreement between two measurement methods, how to read the bias line and limits of agreement on a Bland-Altman plot, how to construct one from a worked dataset, and how to detect proportional bias.
Minimal Clinically Important Difference (MCID): Anchor-Based vs. Distribution-Based Estimation
Compares anchor-based MCID estimation (global rating of change, mean-change and threshold/ROC approaches) against distribution-based estimation (the 0.5 SD rule, SEM-based rules), explains why the two families routinely disagree, and covers which a reviewer will trust more plus how MCID feeds sample-size calculation and responder analysis.
Ecological Validity: Defending It Against the Trade-Off with Internal Validity
Ecological validity means a study’s setting and tasks resemble real-world conditions, and that its findings still hold outside the study. This guide covers what it actually asks, where it sits in the standard validity framework, the concrete design choices (setting, task realism, measurement obtrusiveness, assignment) that trade it off against internal validity, and how to defend it in a methods section.
Test-Retest Reliability: Choosing the Retest Interval, and Checking for Drift
The retest interval is the design decision a test-retest coefficient cannot recover from. What the interval trades off, what COSMIN actually asks about it, which ICC form a test-retest design needs, and why only a mean difference or Bland-Altman check reveals systematic drift between occasions.
Visual Analogue Scale (VAS): Construction, Scoring and Minimal Important Change
A visual analogue scale is a 100 mm line with two anchors and no marks in between, scored by measured distance. This guide covers the construction rules that keep it valid, why VAS and NRS scores are not interchangeable, digital-versus-paper equivalence, and the published minimal important change values with the population, baseline severity and derivation method attached to each.
Inter-Rater Reliability: Choosing the Right Coefficient
A selection table from data type and rater design to the right agreement statistic – Cohen’s and weighted kappa, Fleiss’ kappa, ICC, Krippendorff’s alpha and Gwet’s AC1 – with the kappa paradox worked through three 2×2 tables that share identical 95% agreement.
Average Variance Extracted (AVE): Calculation, the 0.50 Threshold, and Discriminant Validity
AVE is the mean proportion of indicator variance a construct explains rather than error. This guide gives the calculation, the origin and conventional status of the 0.50 threshold, and the Fornell-Larcker and HTMT discriminant-validity tests that consume it, including why the methodological literature has moved away from Fornell-Larcker.
Minimal Detectable Change (MDC): SEM-Based Calculation and the MDC-vs-MCID Decision Rule
What MDC is, the three routes to a standard error of measurement and the assumptions each carries, the MDC95 = 2.77 x SEM calculation worked end to end, the individual-versus-group distinction, and the four-cell decision rule for judging a change against both MDC and MCID.
Cronbach’s Alpha: Reliability & Interpretation Guide
What Cronbach’s alpha does and does not tell you about a scale: the formula and its assumptions, why a high alpha is not evidence of unidimensionality, what the benchmark thresholds are worth, and how to report reliability in APA 7th-edition format.
Criterion Validity: Concurrent and Predictive Evidence Explained
Criterion validity explained: concurrent vs. predictive evidence, how to choose a defensible criterion, criterion contamination, attenuation, and a worked correlation example.
Content Validity: Does Your Instrument Cover the Whole Construct?
Content validity is whether an instrument’s items adequately sample the full construct domain, as judged by expert panels. Covers the expert-panel process, the content validity ratio (CVR), the content validity index (CVI), and the distinction from face validity, each worked through a hand-calculated illustrative example.
Response Bias: The Main Types and How to Design Against Them
Response bias is systematic, direction-consistent error in self-report data. This guide covers the five main types — acquiescence, extreme responding, social desirability, recall bias, and order effects — with the specific design fix for each.
The Hawthorne Effect: Why Being Watched Changes the Data
The Hawthorne effect is the tendency for people to change behavior because they know they are being studied. Learn the Western Electric study history, what the Levitt & List and Jones re-analyses actually found, and how to detect and design against reactivity.
Social Desirability Bias: Why Respondents Tell You What You Want to Hear
What social desirability bias is, the self-deception vs. impression-management mechanisms behind it, where it hits hardest, and the countermeasures (indirect questioning, list experiments, randomized response, self-administration) that reduce it.
Psychometrics: How Researchers Measure Things You Can’t Observe
How psychometrics turns unobservable constructs into reliable, valid scores: the construct-to-score pipeline, classical test theory vs. IRT, and how instruments are built and evaluated.
Levels of Measurement: Nominal, Ordinal, Interval and Ratio
A comparison of the four levels of measurement (nominal, ordinal, interval, ratio), the operations and statistics each one permits, and a step-by-step flowchart for classifying any variable.
Construct Validity: Definition, Evidence Types, and Threats
Construct validity is whether an instrument truly measures the theoretical construct it claims to. This guide covers the Messick/AERA-APA-NCME evidence-argument framing, convergent, discriminant, nomological, known-groups and factorial evidence, the multitrait-multimethod matrix, AVE/HTMT, and the two core threats: construct underrepresentation and construct-irrelevant variance.
The Observer Effect in Research: Reactivity, Not Physics
The observer effect (reactivity) is behavior change caused by awareness of being studied. Learn how it differs from the Hawthorne effect, demand characteristics, social desirability bias, and observer bias, where it bites hardest, and how to mitigate it responsibly.
Accuracy vs Precision in Measurement
Accuracy is closeness to the true value (systematic error); precision is closeness of repeated measurements to each other (random error). How to tell them apart, why “precise but inaccurate” is the dangerous case, and how each connects to reliability, validity, calibration, and diagnostic accuracy.
Reliability in Research: What It Means and How to Assess It
What reliability means in research measurement, how it differs from validity, classical test theory, the four types of reliability, what degrades it, attenuation, and how much reliability is enough.
Intraclass Correlation Coefficient (ICC): Forms, Interpretation, and How to Report It
What the ICC measures, the ICC(1,1)/(2,1)/(3,1) forms and when to use each, Koo & Li interpretation benchmarks, ICC vs. kappa/Pearson r/Bland-Altman/Cronbach’s alpha, sample size, and how to report results.
Types of Validity in Research: Measurement Validity vs Design Validity
Measurement validity (face, content, construct, criterion) and design validity (internal, external, statistical-conclusion) are two distinct families constantly conflated under one word. This guide separates them and covers the validity-reliability relationship.
Cronbach’s Alpha: What It Measures, How to Interpret It, and When to Use Omega Instead
Cronbach’s alpha measures internal consistency, not unidimensionality or reliability in general — the most common misreading of the statistic. This guide covers interpretation, why the 0.7 threshold is misapplied, how adding items inflates alpha, when McDonald’s omega is the better choice, and how to report alpha correctly in a methods section.








