Skip to main content
v2026.11,772 entries · CC-BY 4.0

Bayesian Sample Size Determination

Three genuinely different Bayesian answers to how much data is enough: precision-based estimation, assurance (power averaged over the prior), and sequential Bayes factor stopping rules that remove the fixed-N requirement entirely.

Written and maintained by CASRAI Editorial Board

Last updated

A frequentist power calculation commits to a single number before the study begins: pick one effect size, treat it as if it were known, and solve for the sample size that gives that one scenario 80% (or 90%) power. Bayesian sample size determination replaces that single-point commitment with one of three genuinely different questions: how much data does it take to estimate the effect to a target precision, what sample size gives good power when the effect size itself is uncertain, or — most differently — how would a trial that never fixes N at all actually be run.

Why the fixed-N calculation is a point estimate wearing a formula

A standard power analysis (see CASRAI’s Power Analysis and Sample Size Calculation guide for the frequentist mechanics) asks: for a true effect of exactly this size, at this alpha, what N gives this much power? The number that comes out looks precise, but it is only as good as the single effect-size guess that went in — usually taken from a pilot study, a related published trial, or a minimum clinically important difference the investigators consider worth detecting. If the true effect is smaller than assumed, the trial is underpowered; if larger, resources were spent enrolling more participants than the question needed. Bayesian sample size determination does not eliminate that uncertainty, but it changes what is done with it: instead of pretending the effect size is known, it is treated as a probability distribution from the start, and the sample size calculation is built around that distribution rather than around one point drawn from it.

Precision-based determination: solving for a target interval width

The most direct Bayesian analogue of a frequentist power calculation sets a target not on power but on precision: choose the sample size that gives the posterior credible interval (see Summarizing a Posterior Distribution for how a credible interval is constructed) a target width — for example, a 95% credible interval for a mean difference no wider than a pre-specified clinically meaningful margin. This “Average Length Criterion” approach sidesteps hypothesis testing altogether: the study is not designed to detect a specific effect against a null, it is designed to estimate the effect to a stated level of precision, regardless of which direction it turns out to point. This framing tends to suit descriptive and estimation-focused research questions — prevalence studies, dose-response characterization, biomarker calibration — better than confirmatory hypothesis-testing trials, where a go/no-go decision, not an interval width, is usually the actual deliverable.

Assurance: averaging power over the prior instead of fixing one effect size

For confirmatory work, the more widely used Bayesian approach is assurance, sometimes called Bayesian power. Rather than computing power at one assumed effect size, assurance puts a prior distribution over the plausible effect sizes — often built from a completed pilot study, a meta-analysis of related trials, or elicited expert opinion — and computes the probability that the trial will succeed averaged across that entire distribution, weighted by how plausible each effect size is. The term and the underlying method are most closely associated with a 2005 methodological paper by O’Hagan, Stevens and Campbell in the clinical-trials biostatistics literature, and the practical effect is that assurance is almost always lower than the conventional power figure computed at a single “best guess” effect size, because it also accounts for the real possibility that the true effect is smaller than the best guess. A trial “80% powered” under the conventional calculation might carry an assurance closer to 60-65% once the uncertainty in that best guess is folded in honestly — a gap worth surfacing to a funder or an IRB rather than letting the higher, single-point number stand unqualified.

Sequential Bayes factor design: removing the fixed-N commitment entirely

Both approaches above still end with a single number: enroll N participants, then stop and analyze. Sequential Bayes factor design does not. Instead of pre-committing to a fixed sample size, the study specifies two decision thresholds on the Bayes factor comparing the effect and null hypotheses — for example, stop and declare support for the effect once BF₁₀ exceeds 10 (Jeffreys’ “strong evidence” band), or stop and declare support for the null once BF₁₀ falls below 1/10 — and the data are monitored continuously (or at pre-specified checkpoints) against those two boundaries as participants accrue. The trial stops the moment either threshold is crossed, whether that happens after 40 participants or 400.

This is workable in a way a frequentist sequential design is not without heavy correction, because of a property Bayesian evidence has that a raw p-value does not: the interpretation of a Bayes factor does not depend on the stopping rule that produced it, so checking the data repeatedly and stopping as soon as a threshold is crossed does not inflate a false-positive rate the way repeated significance testing does under a fixed alpha. That does not mean sequential Bayes factor monitoring is free of design choices — the thresholds themselves, and a maximum sample size cap for the (real) scenario where neither boundary is ever crossed, still have to be chosen and justified before data collection starts, exactly the kind of pre-specification a Bayesian adaptive protocol’s statistical analysis plan is expected to document (see Bayesian Adaptive Design for how this fits into a broader adaptive-trial framework, and CASRAI’s Frequentist vs. Bayesian Statistics in Clinical Trials comparison for how the two paradigms differ more generally). Rather than solving the design analytically, the standard workflow — most closely associated with Schönbrodt and Wagenmakers’ Bayes factor design analysis method — simulates the sequential monitoring process thousands of times under both the null and a plausible alternative, and reports, for a given pair of thresholds, the expected sample size at stopping, the probability of reaching a maximum-N cap without a decision, and the long-run rate of each type of error — the same design-characterization exercise a frequentist group sequential design does with alpha-spending functions (see Group Sequential Designs), just built around evidential thresholds instead of a p-value boundary.

Choosing thresholds and a maximum-N cap

Three decisions drive a sequential Bayes factor design in practice:

  • The Bayes factor thresholds themselves. Symmetric thresholds of 10 and 1/10 (Jeffreys’ “strong” band) are a common default; wider gaps (30 or higher) demand more compelling evidence before stopping but generally require a larger expected sample size to get there.
  • The prior on the effect under the alternative hypothesis. Because the Bayes factor averages the likelihood over the prior, a poorly chosen or overconfident alternative prior distorts both the expected sample size and the trial’s actual sensitivity — the same prior-sensitivity concern documented in CASRAI’s Prior Sensitivity Analysis guide applies with full force here, and a design analysis should be re-run under a small set of plausible alternative priors, not just the single one used in the primary plan.
  • A maximum-N cap. Real trials cannot monitor indefinitely; a pre-specified ceiling, chosen from the simulation results so the probability of reaching it without crossing either boundary is acceptably low, keeps the open-ended design operationally and financially bounded, and gives an unambiguous action (typically: stop, report inconclusive) if the cap is reached.

Regulatory considerations

Bayesian sample size approaches are not treated by regulators as inherently riskier than a fixed frequentist calculation, but they are held to a higher documentation standard because their operating characteristics — Type I error rate, expected sample size, power — are usually established through simulation rather than a closed-form formula. FDA’s Adaptive Designs for Clinical Trials of Drugs and Biologics guidance, finalized in the Federal Register in December 2019, expects the full statistical analysis plan for a Bayesian adaptive or sequential design — priors, decision thresholds, and simulation evidence demonstrating the design’s error rates and expected sample size under plausible scenarios — to be pre-specified, justified, and generally discussed with the agency before the trial begins. FDA has separately circulated draft guidance specifically on the use of Bayesian methodology in drug and biological product trials (Federal Register notice, January 2026), and ICH’s E20 guideline on adaptive designs, in development as of 2026, extends the same simulation-and-pre-specification expectations internationally. None of this is unique to sample size determination specifically — it is the general evidentiary bar every Bayesian adaptive or sequential element in a protocol has to clear — but a sample size plan that abandons a fixed N is exactly the kind of design choice that draws that scrutiny first.

A practical workflow

  • Decide what “enough data” means for this study: a target estimation precision, an assurance-adjusted power target, or a pair of evidential stopping thresholds — these are different questions with different deliverables, not three ways of asking the same one.
  • Build the prior (or the prior over effect sizes, for assurance) from a real source — a completed pilot, a meta-analysis, or documented expert elicitation — and report where it came from.
  • Run the design analysis by simulation: expected sample size, probability of hitting a maximum-N cap, and error rates under both the null and the alternative.
  • Re-run the design analysis under at least one alternative, less favorable prior, and report both sets of operating characteristics — not just the primary-plan numbers.
  • Pre-specify the maximum-N cap and the action taken if it is reached, before enrollment opens.
  • Document all of the above in the statistical analysis plan submitted for IRB and (where applicable) regulatory review, in the same level of detail a frequentist group sequential design’s alpha-spending function would require.

Frequently asked questions

Is a Bayesian sample size calculation always smaller than a frequentist one?

Not necessarily, and often the opposite. Assurance folds in genuine uncertainty about the effect size, which typically lowers the achieved probability of success relative to a frequentist calculation anchored to one optimistic point estimate — meaning a larger sample size is often needed to reach the same nominal success probability once that uncertainty is honestly represented, not a smaller one.

Can a sequential Bayes factor design be stopped early just because the result “looks good”?

No — the point of pre-specifying the Bayes factor thresholds before data collection begins is precisely to prevent that. The trial stops only when the accumulating evidence crosses a threshold fixed in advance, and the pre-specification, along with the simulated operating characteristics under that threshold pair, is what a reviewer checks to confirm the decision rule was not adjusted after seeing the data.

Does removing the fixed sample size mean the trial could run forever?

In principle the monitoring is open-ended, which is exactly why a pre-specified maximum-N cap is a required part of the design, not an optional safeguard. The design analysis simulation is what tells investigators how likely the trial is to reach that cap without a decision, so the cap can be set at a sample size that is both operationally feasible and rarely needed in practice.

How does assurance relate to a Bayesian adaptive design?

They address different stages. Assurance is a planning-stage calculation, done before the trial opens, to choose a sample size that accounts for uncertainty in the assumed effect. Bayesian adaptive design describes what happens once the trial is running — using accumulating posterior probabilities to adjust randomization, drop arms, or stop early. A single trial can use an assurance-based sample size at the planning stage and still run as a fixed, non-adaptive design in conduct, or it can combine both.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · included with Regulatory Radar

Ask about Bayesian Sample Size Determination

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.