Skip to main content
v2026.11,772 entries · CC-BY 4.0

Split-Half Reliability and the Spearman-Brown Correction

A split-half correlation is the reliability of a half-length test, not the full one, and different equally valid splits of the same items produce different numbers. What the Spearman-Brown correction fixes, what it can’t fix, and why Cronbach’s alpha — the average of every possible split — became the standard instead.

Written and maintained by CASRAI Editorial Board

Last updated

On this page: why a split-half correlation is only half the answer — the raw correlation between two half-tests describes a test that is itself half as long as the real one, which is why it systematically understates reliability; the Spearman–Brown correction that fixes this; why different, equally defensible ways of splitting the same items into two halves produce different raw correlations, and what that arbitrariness means for how much to trust a single split-half estimate; and how Cronbach’s alpha, introduced four decades later, resolves the arbitrariness by averaging over every possible split rather than committing to one.

Split-half reliability belongs to the same family of internal-consistency methods as Cronbach’s alpha — both estimate reliability from a single administration of a scale, without a second testing occasion. For the across-time counterpart, see CASRAI’s guide to test–retest reliability; for the general vocabulary of reliability types, see reliability in research measurement.

What Split-Half Reliability Actually Estimates

Split-half reliability is computed in three steps. First, the items on a scale are divided into two halves. Second, each respondent’s score is summed separately within each half, producing two scores per person instead of one. Third, the two sets of half-scores are correlated across respondents (typically with a Pearson correlation). That raw correlation, usually written rhh, is treated as an internal-consistency estimate: to the extent the two halves agree on how respondents rank relative to one another, the items appear to be measuring the same thing.

The method has one clear practical advantage over test–retest reliability: it needs only one sitting. There is no retest interval to justify, no regression-to-the-mean confound between occasions, and no risk that respondents change between administrations. What it estimates is narrower in exchange — internal consistency within one occasion, not stability across occasions or agreement across raters. A scale can split cleanly in half and still drift over time, and a scale that is stable over time can still fail a split-half check if its items are heterogeneous. The two properties are independent; see CASRAI’s comparison of test–retest and internal-consistency reliability types for how they diverge in practice.

The Problem the Raw Correlation Creates: A Half-Length Test

The raw half-test correlation cannot be reported as the reliability of the full scale, and the reason is structural, not a matter of rounding. Splitting a k-item scale in half does not just divide the items — it creates two separate, shorter instruments, each with half as many items as the original. Classical test theory treats reliability as a function of test length: holding item quality constant, a longer test is more reliable than a shorter one, because random measurement error attached to any single item averages out further as more items are summed. Correlating two half-length tests against each other therefore produces an estimate of the reliability of a half-length test, not the full one — and it understates the full test’s reliability by a predictable, correctable amount.

This is the specific problem the Spearman–Brown correction exists to solve: not to adjust for a flawed split, but to translate a half-length reliability estimate back into a full-length one.

The Spearman–Brown Correction

The correction is a special case of a more general result published independently and simultaneously by Charles Spearman and William Brown in 1910, in separate papers in the British Journal of Psychology. The general Spearman–Brown prophecy formula predicts how a reliability coefficient changes when a test is lengthened or shortened by a factor n, assuming the added items are equivalent in quality to the existing ones:

ρnn = (n × ρ) / (1 + (n − 1) × ρ)

where ρ is the reliability of the original-length test and ρnn is the predicted reliability after multiplying its length by n. Applied to the split-half case, the test has been shortened to half its length (n = 0.5 relative to the full test, or equivalently, the full test is twice the length of either half, n = 2 going the other direction), which simplifies to the familiar split-half form:

rSB = (2 × rhh) / (1 + rhh)

where rhh is the raw correlation between the two half-scores and rSB is the corrected estimate for the full-length scale. Because the correction is monotonically increasing for any rhh between 0 and 1, rSB is always larger than the raw rhh it is computed from — a raw half-test correlation of 0.70, for example, corrects to an rSB of 0.82. Reporting rhh uncorrected is a common and avoidable error: it is not a conservative or cautious choice, it is simply the wrong number for the question “how reliable is the full scale.”

The general prophecy formula has a second, independent use beyond split-half correction: predicting the effect of lengthening or shortening any scale, which is why it also underlies the standard reasoning for how many items to add to a scale to reach a target reliability, or how much reliability a shortened scale will sacrifice.

The Arbitrary-Split Problem

The correction assumes the two halves are parallel — equivalent in content, difficulty, and the amount of true-score variance they capture. In practice, no single way of splitting a real item set achieves this exactly, and different defensible splits of the identical item pool produce different raw rhh values, which then correct to different rSB values. The three splits most commonly used illustrate why:

  • Odd–even split. Odd-numbered items form one half, even-numbered items the other. This is the conventional default specifically because it distributes items that are adjacent in the test booklet — and therefore similar in content clustering, difficulty progression, and position-related fatigue or practice effects — roughly evenly across both halves, rather than concentrating them in one half.
  • First-half / second-half split. The first k/2 items form one half, the remaining k/2 items the other. This split is vulnerable to any effect that changes systematically across the test’s position — items commonly increase in difficulty toward the end of a scale, respondents fatigue or lose motivation over the course of a longer instrument, and speeded tests in particular produce a first half nearly everyone finishes and a second half many respondents rush or don’t reach. Any of these can depress or inflate rhh for reasons that have nothing to do with the items’ actual internal consistency.
  • Random split. Items are assigned to one half or the other by a random draw rather than a fixed rule. This avoids any systematic content or position confound but is not repeatable — a different random draw on the identical data produces a different rhh, and with a modest number of items, the number of distinct possible splits is large enough that this variability is not a rounding-level concern.

The consequence is that “the split-half reliability of this scale” is not a single, well-defined number the way a test–retest coefficient or a Cronbach’s alpha is — it depends on a split-assignment decision that the analyst makes, and that decision is not fully specified by the method itself. Two analysts working from the same dataset can report two different, both procedurally correct, split-half reliabilities for the same scale. For a heterogeneous item set (items covering somewhat different facets of a construct, or items with a wide spread of difficulty), the gap between the best-case and worst-case split can be substantial; for a short, homogeneous item set it is usually small. Reporting a split-half estimate without stating exactly how the split was formed is therefore incomplete in a way that matters — a reader cannot tell whether the reported number is close to the average outcome across possible splits or came from whichever split happened to be tried.

A Worked Example

Consider a 10-item scale administered to a sample of respondents. Summing the odd-numbered items (1, 3, 5, 7, 9) into one half-score and the even-numbered items (2, 4, 6, 8, 10) into the other, and correlating the two half-scores across respondents, might yield rhh = 0.65. Applying the correction: rSB = (2 × 0.65) / (1 + 0.65) = 1.30 / 1.65 ≈ 0.79. If instead the first five items are compared against the last five, and item 9 and item 10 happen to be the two items most respondents rush through because they run out of time, the raw correlation for that split could come out lower — say rhh = 0.52, correcting to rSB ≈ 0.68. Both numbers were computed correctly from the identical dataset; they differ because they answer a slightly different question about a slightly different pair of half-tests. Neither is “the” split-half reliability of the scale — each is the split-half reliability of one particular split of it.

From One Split to Every Split: How Cronbach’s Alpha Supersedes Split-Half

Lee Cronbach’s 1951 Psychometrika paper introducing coefficient alpha did not propose a competing method so much as resolve the arbitrary-split problem directly. For a k-item scale, there are many distinct ways to divide the items into two equal halves. Cronbach’s alpha is mathematically equivalent to computing the Spearman–Brown-corrected split-half reliability for every possible way of splitting the scale in half, and averaging the results. Because it is built from the average inter-item covariance rather than from any one arbitrary partition, alpha does not depend on a split-assignment decision at all — it produces the single number that a split-half analysis, taken to its logical limit across every possible split, converges on. This is the direct sense in which alpha supersedes split-half reliability: it is not a different concept competing for the same purpose, it is what split-half reliability becomes once the arbitrariness of picking one split is removed.

There is an older, narrower predecessor worth distinguishing here too: the Kuder–Richardson formulas (1937), specifically KR-20, compute the same all-possible-splits average as alpha but only for dichotomously scored (right/wrong, yes/no) items. Alpha is the general form that also handles Likert-type and other continuously scored items; KR-20 is mathematically a special case of alpha for binary items, not a separate method.

None of this means a split-half estimate is now the wrong thing to compute in every context — see the next section — but for the ordinary purpose of reporting the internal consistency of a multi-item scale, alpha (or, where alpha’s own assumptions are in doubt, McDonald’s omega) has been the standard choice for decades precisely because it removes the split-selection decision from the analyst’s hands. CASRAI’s guide to Cronbach’s alpha covers what the statistic assumes, how it is commonly misinterpreted, and when omega is the better choice; the SPSS mechanics of computing it, including the item-total statistics that identify a weak item, are covered separately in the guide to running Reliability Analysis in SPSS.

When Split-Half Reliability Is Still Worth Computing

Split-half reliability is rarely the first choice for reporting a scale’s internal consistency today, but it has not disappeared, for a few concrete reasons:

  • Historical and legacy instruments. Older published scales, and some legacy clinical or educational instruments, report split-half coefficients from the era before alpha was in routine use. Understanding what that number does and doesn’t mean is necessary to compare a historical figure against a modern alpha reported for the same or a revised instrument — they are not interchangeable without the caveats above.
  • Speeded tests. Alpha and KR-20 assume every respondent attempts every item; on a strictly speeded test, where the score mostly reflects how many items a respondent reached rather than how many they answered correctly, alpha is known to overestimate reliability because unreached items at the end correlate near-perfectly with each other (everyone scores them the same way — unattempted). A split-half approach using an odd–even split, applied only to items every respondent actually reached, avoids this specific distortion in a way alpha does not by default.
  • Teaching and illustration. Because it is computed from an ordinary correlation most students already understand, split-half reliability remains a common way to introduce the concept of internal-consistency reliability before introducing alpha’s covariance-based formula.
  • A quick sanity check. An odd–even split-half estimate, computed alongside alpha, is sometimes reported as a secondary cross-check — if the two diverge sharply, that is itself informative about how heterogeneous the item set is.

Reporting a Split-Half Estimate Correctly

Where a split-half figure is reported at all, methodological guidance converges on the same minimum: state the split method used (odd–even is the default that should be assumed only if stated explicitly — don’t leave a reader to guess), report the raw rhh and the corrected rSB separately rather than only the corrected figure, and report the sample size the correlation was computed on. A split-half coefficient without the split method named is not fully interpretable, for exactly the reason set out above — the same dataset can produce a materially different number under a different, equally legitimate split.

Frequently Asked Questions

Is split-half reliability the same as internal consistency?

Split-half reliability is one method of estimating internal consistency, not a synonym for the broader concept. Cronbach’s alpha and McDonald’s omega are the other common internal-consistency estimates computed from a single administration; all three ask the same underlying question (do the items on this scale correlate with each other the way items measuring one construct should?) using different computational approaches, and alpha’s all-splits average is generally preferred today for the reasons above.

Why does the Spearman–Brown-corrected value always come out higher than the raw correlation?

Because the raw correlation is the reliability of a test half as long as the real one, and reliability increases with test length under classical test theory (more items means random item-level error averages out further). The correction translates the half-length estimate back up to the full-length test’s predicted reliability, and that translation is mathematically guaranteed to increase the number for any positive raw correlation.

What counts as an acceptable split-half or Spearman–Brown-corrected reliability value?

The same conventional thresholds usually applied to alpha are typically applied here — roughly 0.70 as a minimum for research use and 0.90+ expected for individual-level clinical or high-stakes decisions — but because the raw figure depends on which split was used, a threshold judgment on a split-half value is only as trustworthy as the split it came from. A borderline result on one split and a comfortable pass on another is a sign to check alpha rather than to pick the more favorable split.

Does the Spearman–Brown formula only apply to split-half reliability?

No — the general prophecy formula predicts the effect of lengthening or shortening a test by any factor, not just doubling a half-length test back to full length. It is the same formula used, in its general form, to estimate how much reliability a shortened scale will lose, or how many additional equivalent items would be needed to reach a target reliability.

Can split-half reliability be computed for a scale with an odd number of items?

Yes, with an uneven split (for example 5 items in one half and 6 in the other on an 11-item scale), though this adds yet another source of variation between possible splits on top of content and position, and most treatments recommend an even-item-count scale, or dropping to the nearest even split deliberately and reporting that it was done, rather than leaving the imbalance unstated.

Sources

  • Spearman C. Correlation calculated from faulty data. British Journal of Psychology. 1910;3(3):271–295.
  • Brown W. Some experimental results in the correlation of mental abilities. British Journal of Psychology. 1910;3(3):296–322.
  • Cronbach LJ. Coefficient alpha and the internal structure of tests. Psychometrika. 1951;16(3):297–334.
  • Kuder GF, Richardson MW. The theory of the estimation of test reliability. Psychometrika. 1937;2(3):151–160.
  • Nunnally JC, Bernstein IH. Psychometric Theory. 3rd ed. McGraw-Hill; 1994. Standard reference for the classical test theory relationship between test length and reliability underlying the Spearman–Brown formula.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Split-Half Reliability and the Spearman-Brown Correction

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.