Written and maintained by CASRAI Editorial Board
Last updated
A sampling frame is the actual list, register, or procedure used to select a sample — and coverage error is the gap between that frame and the target population it is supposed to represent. Every probability sampling method that operates downstream of frame selection (simple random, stratified, cluster, systematic) assumes the frame is a workable stand-in for the population. When it isn’t, correct randomization inside the frame does not fix what the frame excluded, duplicated, or wrongly included before sampling even started.
This page is about the frame itself: what makes a frame diverge from its target population, the two directions that divergence takes, how to build and assess one, and — because this is the part most methods sections skip — exactly what a reader needs you to report about the frame to judge whether your results generalize.
What Coverage Error Is, and What It Isn’t
Three things are easy to conflate and worth separating precisely:
- Target population — the group you want to draw conclusions about.
- Sampling frame — the concrete list, register, or procedure you actually use to identify and select units.
- Sample — the subset of the frame you select and, ideally, obtain data from.
Coverage error is the mismatch between the first two — it is fixed the moment the frame is built, before a single unit is drawn. This is what distinguishes it from the other error sources a reader might lump it in with:
- Sampling error exists even with a perfect frame, purely because you observe a sample rather than a census; it shrinks as sample size grows.
- Non-response error occurs after selection, when sampled units fail to respond.
- Coverage error occurs before selection, in how the frame itself was built, and a larger sample does not fix it — it just estimates the frame’s population more precisely, not the target population’s.
This separation comes from the Total Survey Error framework laid out in Groves, Fowler, Couper, Lepkowski, Singer, and Tourangeau’s Survey Methodology (Wiley), the standard reference that splits survey error into representation errors (coverage, sampling, non-response, adjustment) and measurement errors (validity, measurement, processing). Coverage error sits on the representation side, and it is the only one of the four representation errors that is entirely determined before sampling begins.
Undercoverage and Overcoverage: the Two Failure Directions
Undercoverage means the frame omits members of the target population who had no chance of selection at all. Common patterns: a landline random-digit-dial frame that omits cell-only households; an alumni email list that omits graduates whose contact record went stale; an employee roster pulled from HR systems that omits staff hired after the last data extract; a property-tax roll used as a household frame that omits renters and unlisted units.
Overcoverage means the frame includes units that are not actually in the target population, or includes the same eligible unit more than once. Common patterns: a merged mailing list that lists one person under two addresses; a customer database that retains records for accounts that closed before the study period; a directory that has not been purged of departed or deceased individuals, giving them a nonzero (and wrong) chance of selection.
The two are not mutually exclusive. A frame assembled from several administrative sources routinely does both at once — missing an entire subgroup while also double-listing part of the population it does cover. Where you can obtain an independent count of the eligible population (a census figure, an administrative registry total), a simple coverage rate is eligible frame units divided by target-population units; where you cannot, the honest move is a qualitative statement of what is known to be excluded, not a fabricated rate.
Building a Frame: Sourcing and Assessing It
Frame construction goes in this order, not in reverse:
- Define the target population precisely first. A frame can only be judged against a population definition that already exists — decide the eligibility rules (who, where, when) before evaluating any candidate frame source. A vague population definition makes every coverage judgment downstream unfalsifiable; see generalizability in research for how that definition feeds directly into what you can later claim your results apply to.
- Identify candidate frame sources. An existing administrative list or register, a membership or professional directory, a geographic/area frame built from addresses or parcels, or a frame you construct yourself when no adequate list exists.
- Assess each candidate against the population definition for known exclusions (who systematically can’t be on this list) and known duplication risk (who might appear more than once) before committing to it — not after data collection is already underway.
- Where no single source adequately covers the population, consider combining frames rather than accepting a large, undocumented gap in one.
When One Frame Isn’t Enough: Multi-Frame Designs
Multi-frame (or dual-frame) sampling combines two or more frames with different, overlapping coverage of the population — the classic case is a landline RDD frame paired with a cell-phone frame, so that cell-only households (undercovered by the landline frame alone) get a nonzero chance of selection. Combining frames without double-counting units that appear on more than one of them requires a compositing estimator that accounts for each unit’s actual probability of selection across all frames it belongs to. That weighting mechanics is substantial enough to be its own topic; the design decision itself — that you used more than one frame, and why — belongs in your methods section regardless of how the weighting was handled.
What Your Methods Section Must Report About the Frame
AAPOR’s Code of Professional Ethics and Practices, Standards for Disclosure (Section III, current as of the April 2021 revision) sets the actual disclosure bar here, and it applies to any publicly released survey result, not just AAPOR-member work. It requires researchers to be explicit about whether the sample comes from a probability-based frame or was selected using non-probability methods; if a frame, list, or panel was used, to name the supplier and describe the nature of the list; to state the decision rules used to define the study population (location, time, and relevant demographic bounds); and — the part most directly about coverage — to describe “the coverage of the population, including describing any segment of the target population that is not covered by the design.”
In practice, that resolves into six things a methods section should state about the frame:
- Source and vintage — what the frame is and when it was current (e.g., “institutional roster extracted March 2025”), not just “a list of employees.”
- Probability or non-probability — stated explicitly, not left for the reader to infer from the sampling method described later.
- Known exclusions — which segments of the target population the frame is known not to reach, and why.
- Duplication handling — whether the frame was deduplicated, and how, if overcoverage was a realistic risk.
- Coverage rate, if calculable — or an explicit qualitative statement in its place if no independent population count exists to compute one against.
- If multiple frames were combined — that fact, stated plainly, even if the compositing/weighting detail is deferred to an appendix.
Note this is a real methodological standard already familiar from AAPOR’s Standard Definitions (10th edition, 2023), which is itself organized by frame rather than by survey mode — a structural choice that reflects how central frame decisions are to everything reported afterward, including response-rate calculations.
Coverage Error Against the Rest of Total Survey Error
It is worth keeping the distinctions straight when you write the limitations paragraph, because each error source implies a different fix. Sampling error shrinks with a bigger sample. Coverage error does not — a bigger sample drawn from an uncorrected bad frame just estimates the frame’s population more precisely, not the target population’s, which is why sampling bias introduced at the frame stage survives any amount of additional data collection. Non-response error is a separate, later-stage problem with its own measurement procedures. And once coverage, sampling, and non-response error are all accounted for, what remains is the generalizability question proper — see internal vs. external validity and generalizability in research for how a well-documented frame feeds directly into what a reader is entitled to conclude from your sample.
Frequently Asked Questions
What’s the difference between a sampling frame and a target population?
The target population is who you want to draw conclusions about. The sampling frame is the actual, concrete list or procedure you use to identify and select units. They are rarely identical; the size of that gap is coverage error.
What is coverage error, in one sentence?
The error introduced by a mismatch between a sampling frame and the target population it is meant to represent — through omission (undercoverage), wrongful inclusion or duplication (overcoverage), or both.
Does increasing sample size fix undercoverage or overcoverage?
No. A larger sample drawn from the same flawed frame produces a more precise estimate of the frame’s population, not the target population. Fixing coverage error requires fixing or supplementing the frame, not sampling more from it.
How do I calculate a coverage rate?
Eligible frame units divided by target-population units, when an independent count of the target population exists (e.g., a census or administrative total) to compare against. Without that independent count, report the known exclusions qualitatively rather than computing a number you can’t actually support.
What must I report about my sampling frame in a methods section?
At minimum, per AAPOR’s Standards for Disclosure: the frame’s source and vintage, whether it is probability or non-probability, the segments of the target population it’s known not to cover, how duplication was handled if relevant, a coverage rate where calculable, and whether multiple frames were combined.








