Skip to main content
v2026.11,772 entries · CC-BY 4.0

Cohort Study: Design, Types, and How It Works

A guide to cohort study design in clinical and epidemiological research: prospective vs. retrospective cohorts, how cohort studies compare to case-control, cross-sectional, and RCT designs, key measures (incidence, relative risk, hazard ratio), and their core strengths and limitations.

Written and maintained by CASRAI Editorial Board

Last updated

A cohort study is an observational design in which researchers group participants by their exposure status — whether or not they have a characteristic, behavior, or exposure of interest — and follow those groups forward in time to see whether and how an outcome develops. Because the exposure is identified before the outcome occurs, cohort studies establish a stronger temporal sequence than case-control or cross-sectional designs, which is why they remain a primary tool for estimating incidence and studying how exposures relate to later health outcomes.

This guide covers what a cohort study is, how it differs from the other major observational designs, the two main ways cohorts are constructed (prospective and retrospective), the strengths and limitations that determine when a cohort design is the right choice, and how cohort findings are typically reported and interpreted.

What makes a study a cohort study

Three features define the cohort design:

  • Grouping by exposure, not outcome. Participants are classified into an exposed group and an unexposed (or comparison) group based on a characteristic or exposure they already have or plan to undergo — not based on whether they’ve experienced the outcome under study.
  • Forward-looking (longitudinal) follow-up. The cohort is tracked over a defined period, either as events happen in real time or by reconstructing that period from existing records, to observe who develops the outcome and who doesn’t.
  • No investigator-assigned intervention. The investigator observes and measures a naturally occurring exposure rather than assigning it. This is what distinguishes a cohort study from a randomized controlled trial (RCT): in an RCT, the researcher decides who receives which intervention; in a cohort study, the researcher measures an exposure that already exists in the population, such as smoking status, occupational history, or a genetic marker.

According to the STROBE Statement (STrengthening the Reporting of OBservational studies in Epidemiology), cohort, case-control, and cross-sectional studies are the three main analytical designs used in observational research. Each answers a different kind of question and carries different trade-offs, discussed below.

This is the dynamic behind the IMPACC long COVID cohort.

Prospective vs. retrospective cohort studies

Cohort studies come in two timing variants, distinguished by when the outcome data is collected relative to when the study begins:

  • Prospective cohort study. The investigator identifies exposure groups now and follows them forward, collecting outcome data as it occurs. This is the classic cohort design — exposure is measured before the outcome exists, which minimizes recall bias and allows the investigator to control data quality throughout follow-up. The trade-off is time and cost: a prospective cohort can take years or decades to accumulate enough outcome events, particularly for rare or slow-developing conditions.
  • Retrospective (historical) cohort study. The investigator uses existing records — medical charts, registries, employment records, insurance claims — to reconstruct exposure status and outcomes that have already occurred. The logical structure is identical to a prospective cohort (group by exposure, look forward from that exposure point to the outcome); only the data source and timing of the researcher’s involvement differ. Retrospective cohorts are faster and cheaper because the follow-up period already happened, but they depend on the completeness and accuracy of existing records, which the investigator did not control at the time they were created.

See the companion comparison, Prospective vs. Retrospective Study, for a fuller treatment of how this timing distinction affects bias and evidence strength across study types generally, not just cohorts.

Cohort studies vs. case-control studies vs. cross-sectional studies vs. RCTs

Design How participants are selected Direction of inquiry Best suited for Main limitation
Cohort study Grouped by exposure status Forward from exposure to outcome Common outcomes; estimating incidence; studying multiple outcomes from one exposure Inefficient for rare outcomes; loss to follow-up over time
Case-control study Grouped by outcome status (cases vs. controls) Backward from outcome to exposure Rare outcomes; generating hypotheses about multiple exposures at once Recall bias; difficulty selecting a truly comparable control group
Cross-sectional study Neither — a defined population sampled at one point in time Exposure and outcome measured simultaneously Estimating prevalence; quick, low-cost snapshots Cannot establish which came first, exposure or outcome
Randomized controlled trial (RCT) Randomly assigned to intervention or comparison arm Investigator assigns exposure, then observes forward Establishing causal effect of an intervention Cost, time, ethical limits on what can be randomly assigned

The choice between cohort and case-control design in particular comes down to how common the outcome is and how much time and budget the study has. Because a cohort study has to enroll enough people to observe a meaningful number of outcome events, it becomes impractical for rare diseases — a case-control study, which starts from people who already have the outcome, is far more efficient there. Conversely, when a single exposure might plausibly cause several different outcomes, a cohort design lets researchers observe all of them in the same followed population, which a case-control study — built around one outcome at a time — cannot do as directly.

For the broader design landscape this guide sits within, including where quasi-experimental and adaptive designs fit, see Clinical Study Design: The Major Types and How They Relate. For the observational-vs-experimental distinction specifically, see RCT vs. Observational Study.

Why cohort studies matter: what they can (and can’t) show

Cohort studies are one of the few observational designs that can estimate incidence — the rate at which new cases of an outcome develop in a population over time — because the study follows people who don’t yet have the outcome and counts how many develop it. Case-control and cross-sectional designs generally cannot do this directly, since they don’t track a defined population forward from a starting point.

Because exposure status is established before the outcome is known, cohort studies also support a stronger argument that exposure preceded outcome than either case-control or cross-sectional designs allow — a necessary, though not sufficient, condition for a causal claim. Cohort studies are, however, still observational: the investigator did not randomly assign who was exposed, so systematic differences between the exposed and unexposed groups (confounding) can distort the association observed. Only randomization, as used in an RCT, reliably balances both known and unknown confounders between comparison groups before the intervention begins. Cohort studies address confounding after the fact, through study design choices (matching, restriction) and statistical adjustment (stratification, multivariable regression), which reduce but do not eliminate the risk that an observed association reflects something other than a true causal effect.

Common measures reported from cohort studies

  • Incidence — the proportion or rate of a study population that develops the outcome over the follow-up period, calculated separately for the exposed and unexposed groups.
  • Relative risk (risk ratio) — the incidence in the exposed group divided by the incidence in the unexposed group. A relative risk of 2.0 means the exposed group developed the outcome at twice the rate of the unexposed group over the same follow-up period.
  • Attributable risk — the absolute difference in incidence between the exposed and unexposed groups, used to estimate how much of the outcome’s occurrence in the exposed group is attributable to the exposure itself, as distinct from the baseline rate that would have occurred anyway.
  • Hazard ratio — used when follow-up time varies across participants (common in long-running cohorts with staggered enrollment or losses to follow-up), estimated through survival-analysis methods such as Cox proportional hazards regression rather than a simple ratio of two incidence proportions.

These are distinct from the odds ratio typically reported in case-control studies, where the outcome-based sampling means true incidence in the source population usually can’t be calculated directly.

Strengths of the cohort design

  • Can measure incidence and multiple outcomes from a single exposure.
  • Establishes exposure-before-outcome timing more convincingly than case-control or cross-sectional designs.
  • Reduces recall bias for exposure data when conducted prospectively, since exposure is recorded before the outcome exists and can’t be colored by knowledge of who later developed it.
  • Well suited to studying rare exposures, since the design can specifically enroll or over-sample people with an uncommon exposure and compare them to an unexposed group.

Limitations of the cohort design

  • Inefficient for rare outcomes. Studying a condition that occurs in a small fraction of the population requires enrolling very large numbers of participants, following them for a long time, or both — often making a cohort design impractical compared with a case-control study for a rare disease.
  • Loss to follow-up. Participants move, withdraw, or become unreachable over long follow-up periods. If those who are lost differ systematically from those who remain — a pattern often called attrition bias — the results can be distorted, particularly if loss to follow-up is related to both exposure and outcome.
  • Cost and duration. Prospective cohorts, in particular, can take years or decades to accumulate sufficient outcome events, requiring sustained funding, participant retention efforts, and infrastructure for long-term data collection.
  • Confounding. Because exposure isn’t randomly assigned, the exposed and unexposed groups may differ in other ways that also affect the outcome. A frequently cited example is the “healthy volunteer” or “healthy worker” effect, where people who choose or are able to maintain a given exposure (such as continued employment) tend to be healthier at baseline than the comparison group, which can bias the observed association toward showing the exposure as more protective than it actually is.
  • Record quality, for retrospective cohorts. When exposure and outcome data are drawn from records created for another purpose (clinical charts, claims data, registries), the investigator has no control over how consistently or accurately that data was originally recorded.

How cohort study results are typically reported

The STROBE Statement provides the widely used reporting checklist for observational studies, including a cohort-specific set of items covering how the cohort was assembled, the sources and methods used to ascertain exposure and outcome, how loss to follow-up was handled, and how confounding was addressed in the analysis. Journals in epidemiology, public health, and clinical medicine commonly require or recommend STROBE-compliant reporting for submitted cohort studies, and research administrators supporting investigators through submission should expect a completed STROBE checklist to be requested as part of the manuscript package for an observational study, similar to how CONSORT applies to randomized trials. See the CONSORT Statement guide for the RCT-side equivalent of this reporting expectation.

IRB, consent, and data-management considerations for cohort studies

Because cohort studies often run for years or decades and depend on repeatedly locating and re-contacting the same people (or reusing existing records at scale), they raise oversight and data-management questions that a shorter, single-visit study does not. Research administrators supporting a cohort protocol typically need to plan for the following, in addition to the standard human-subjects review any study requires.

  • Continuing review over a long follow-up period. Under the 2018-revised Common Rule, certain minimal-risk research is exempt from mandatory continuing review (45 CFR 46.109(e)), but many multi-year prospective cohorts do not meet that exemption — either because risk increases as new sub-studies or biospecimen uses are added, or because an institution’s own policy requires periodic re-review regardless. A cohort protocol’s IRB submission should specify the review interval up front, since re-approval lapses can halt enrollment or follow-up contact.
  • Consent for prospective cohorts. Standard informed consent is obtained at enrollment, but because investigators frequently cannot fully specify every future analysis or biospecimen use decades in advance, many prospective cohorts now use broad consent — a single consent, formalized as a distinct regulatory option by the 2018 Common Rule revision (45 CFR 46.116(d)), covering storage and future secondary research use of identifiable data or biospecimens. Broad consent has its own required elements and does not cover every possible future use without limit; where it was declined or wasn’t offered, a new consent or a separate IRB determination is needed before reusing a participant’s data or specimens in a new sub-study.
  • Waiver of consent for retrospective cohorts. A retrospective cohort built from existing charts, claims, or registry data commonly cannot obtain fresh consent from everyone whose records are used. The IRB pathway for this is a waiver of the consent process itself under 45 CFR 46.116(f), which requires the IRB affirmatively find: the research involves no more than minimal risk; the waiver won’t adversely affect participants’ rights and welfare; the research could not practicably be carried out without it; and, where appropriate, participants will be debriefed. This is an IRB determination, not something an investigator can self-certify, and it is a different, narrower provision than the 46.117(c) waiver of the signed consent form (which still requires a consent process, just not necessarily a signature).
  • HIPAA runs in parallel to, not instead of, the Common Rule. When a retrospective cohort draws on records held by a HIPAA covered entity, using protected health information without individual authorization requires a separate waiver from the IRB or a Privacy Board under 45 CFR 164.512(i) — a distinct legal analysis from the 46.116(f) waiver, with its own criteria. Getting one waiver does not automatically satisfy the other; both may be needed for the same retrospective cohort.
  • Data management planning for decades-long follow-up. A cohort’s data management plan needs to account for things a shorter study’s DMP typically does not: how contact and locating information is kept current and secured separately from research data, how data collected across many follow-up waves is linked to the same participant over time without compromising confidentiality, what happens to already-collected data if a participant withdraws partway through follow-up (data already collected is typically retained unless the participant specifically requests destruction and that is feasible), and a retention schedule that may need to extend well beyond a typical grant’s post-award retention period because the cohort itself, or secondary analyses of it, may continue for years after the original funding ends.

See CASRAI’s Data Management Plan (DMP) entry and Informed Consent in Research guide for the underlying mechanics of each of these pieces.

Frequently asked questions

Is a cohort study the same as a clinical trial?

No. A cohort study is observational — the investigator measures an exposure that already exists in the population rather than assigning it. A clinical trial, and specifically a randomized controlled trial, is interventional: participants are prospectively assigned to receive a specific intervention according to a protocol. Some large cohort studies do incorporate embedded trials or sub-studies, but the base cohort design itself is not a trial.

What’s the difference between a cohort study and a longitudinal study?

“Longitudinal” describes any study that follows the same participants over time, which includes cohort studies but also other designs (some registries, certain panel surveys) that aren’t organized around comparing exposed and unexposed groups. A cohort study is a specific type of longitudinal study defined by the exposure-based grouping described above.

Can a cohort study prove causation?

Not on its own. A well-designed cohort study can establish that exposure preceded outcome and can adjust for known confounders, which strengthens the case for causality relative to case-control or cross-sectional evidence — but because exposure isn’t randomly assigned, unmeasured or residual confounding can never be fully ruled out. Causal inference from cohort data typically depends on converging evidence: consistency across multiple studies, a plausible biological or mechanistic pathway, and often triangulation with trial evidence where an RCT is feasible.

How large does a cohort need to be?

It depends entirely on how common the outcome is and how large an effect the study is designed to detect. A formal sample size calculation, based on the expected incidence in the unexposed group and the smallest relative risk the study needs to reliably detect, is standard practice before a cohort study begins enrollment; a biostatistician is typically consulted for this step during protocol development.

What’s an example of a well-known cohort study?

Large, long-running cohort studies such as the Framingham Heart Study (which has followed participants and their descendants since 1948 to study cardiovascular disease) and the Nurses’ Health Study are frequently cited examples of how sustained cohort follow-up has shaped understanding of chronic disease risk factors over time. These illustrate both the strength of the design — decades of exposure and outcome data on the same population — and its resource demands.

Does a cohort study need IRB approval?

Yes, if it involves human subjects as defined by the Common Rule (45 CFR 46) or FDA human-subjects regulations — which covers essentially every cohort study using identifiable data or biospecimens from living people, whether prospectively collected or drawn from existing records. A prospective cohort typically requires standard informed consent (often supplemented with broad consent for future secondary use); a retrospective cohort built from existing records commonly proceeds under an IRB waiver of consent rather than fresh consent from every person in the dataset. Either way, the determination of which pathway applies, and whether the study qualifies for exempt, expedited, or full-board review, is made by the IRB, not the investigator.

Related reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Cohort Study: Design, Types, and How It Works

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.