Skip to main content
v2026.11,772 entries · CC-BY 4.0

Direct comparison

Test-Retest Reliability vs. Inter-Rater

Test-retest reliability = the same measure given to the same sample twice, then correlated (Pearson r or ICC), usually 2-4 weeks apart. Vs. inter-rater.

Written and maintained by CASRAI Editorial Board

Last updated

Ask CASRAI · included with Regulatory Radar

Ask about Test-Retest Reliability vs. Inter-Rater

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

How do Test-Retest Reliability, Inter-Rater Reliability compare side by side?

The table below compares Test-Retest Reliability, Inter-Rater Reliability across 11 procurement-relevant dimensions, from what it measures through typical interval between comparisons.

Side-by-side comparison

DimensionTest-Retest ReliabilityInter-Rater Reliability
What it measuresConsistency of the same measure across timeConsistency of the same measure across different raters or observers
What varies between the two measurementsTime (same rater/instrument, two occasions)Rater or observer (same or near-same occasion, two or more raters)
What is held constantRater and instrumentTiming and the material/target being rated
Statistic for continuous dataPearson's r or intraclass correlation coefficient (ICC)Intraclass correlation coefficient (ICC)
Statistic for categorical dataCohen's kappa (same rater, two time points)Cohen's kappa (2 raters) or Fleiss' kappa (3+ raters); raw percent agreement is a weaker, chance-uncorrected alternative
Common interpretation benchmarkICC: <0.50 poor, 0.50-0.75 moderate, 0.75-0.90 good, >0.90 excellent (Koo & Li, 2016)Same ICC bands for continuous ratings; kappa: <0 poor, 0-0.20 slight, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, 0.81-1.00 almost perfect (Landis & Koch, 1977)
Main threat to a good resultPractice/memory effects, true change in the construct over the interval, regression to the meanAmbiguous coding criteria, rater bias or leniency differences, insufficient rater training/calibration
Number of raters/instruments requiredOne (same rater or instrument, two administrations)Two or more
Typical use caseValidating a survey, questionnaire, diagnostic instrument, or physiological measure meant to be stable over timeValidating a coding scheme, clinical diagnosis, qualitative theme coding, or systematic-review screening decisions
Often confused withInternal consistency (Cronbach's alpha) -- that measures item homogeneity within a single administration, not stability over timeIntra-rater reliability -- the same single rater scoring the same material twice, not different raters
Typical interval between comparisonsLong enough that respondents cannot simply recall earlier answers, short enough that the construct hasn't genuinely changed -- often 2-4 weeksNot time-based -- raters typically score the same fixed set of cases independently, close together in time

Common questions

Common questions about Test-Retest Reliability vs Inter-Rater Reliability

What is the basic difference between test-retest and inter-rater reliability?

+

Test-retest reliability holds the rater and instrument constant and varies time -- it asks whether the same measurement process gives the same answer twice. Inter-rater reliability holds timing constant and varies the rater -- it asks whether different people scoring the same material agree with each other.

Can a measure have high test-retest reliability but low inter-rater reliability?

+

Yes. A single well-trained rater can be highly consistent with themselves over time while a second rater, working from an ambiguous or under-specified coding scheme, produces very different scores. The two statistics isolate different sources of measurement error, so neither one guarantees the other.

What counts as a good ICC value for reliability?

+

The commonly cited Koo & Li (2016) guideline treats ICC values below 0.50 as poor, 0.50-0.75 as moderate, 0.75-0.90 as good, and above 0.90 as excellent reliability. The same bands are typically applied to both test-retest and inter-rater ICCs, though acceptable thresholds vary somewhat by field and by how the measure will be used.

Is Cohen's kappa the same thing as inter-rater reliability?

+

No -- Cohen's kappa is one specific statistic used to quantify inter-rater reliability for categorical data with two raters. Inter-rater reliability is the broader concept; the appropriate statistic depends on the data type and number of raters (ICC for continuous ratings, Fleiss' kappa for three or more raters on categorical data, Cohen's kappa for exactly two raters on categorical data).

How does test-retest reliability differ from internal consistency (Cronbach's alpha)?

+

Test-retest reliability assesses stability across two separate administrations of the same measure over time. Internal consistency (commonly reported as Cronbach's alpha) assesses whether items within a single administration of a multi-item scale correlate with each other, i.e. whether they appear to measure the same underlying construct. A scale can have high internal consistency in one sitting and still show poor test-retest stability, or vice versa.

What is a typical time interval for a test-retest reliability study?

+

There is no universal standard, but two to four weeks is a common default in survey and psychometric validation research -- long enough to reduce the chance that respondents simply recall their earlier answers, short enough that the underlying trait or attitude being measured is unlikely to have genuinely changed.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.