Abstract
This article traces the historical development of frequentist sample size estimation from its philosophical origins to its present-day complexity. Preliminary concepts were identified by Christiaan Huygens’ work on expected value and Jacob Bernoulli’s law of large numbers, which first linked sample size and estimation accuracy. The eighteenth and nineteenth centuries brought major advances in probability theory through the work of Pierre-Simon de Laplace, Carl Friedrich Gauss, and Siméon-Denis Poisson, yet explicit sample size planning remained uncommon. The early twentieth century saw the emergence of methods for sample size calculations based on Jerzy Neyman and Egon Pearson’s hypothesis testing framework and Sir Ronald Aylmer Fisher’s experimental design principles. While Donald Mainland and Austin Bradford Hill referred indirectly to these as early as the 1930s, it took many decades before their explicit use became common. After the Second World War, contributions from figures such as Abraham Wald further embedded sample size planning with sequential methodologies. From the 1970s onward, standardised formulas, regulatory requirements, reporting standards such as CONSORT and statistical software consolidated frequentist sample size estimation as a routine component in applied research. In the twenty-first century, simulation-based, adaptive and Bayesian approaches, together with open-source computational ecosystems, have expanded the scope and accessibility of sample size methods. In contemporary research, sample size estimation has evolved into a multifaceted discipline; its methodological sophistication is contingent upon the underlying objective, illustrating the persistent divergence between explanatory inference and decision-oriented design.
Introduction
In the realm of empirical research, determining the appropriate sample size is a critical component of study design. Whether estimating a population proportion, testing a new treatment, or validating a machine learning model, researchers must decide how much data to collect to ensure results are reliable, reproducible, and scientifically meaningful.
Sample size estimation is not just a statistical technicality; it reflects broader epistemological commitments. Collect too little data, and the study’s findings lack statistical power, risking false negative or false positive inferences. Collecting too much may waste resources or expose more participants than necessary to experimental or control risks. Finding the sample size that is ‘just right’ is both an art and a science.
Historically, the development of sample size estimation has mirrored the evolution of probability and statistics, from early philosophical inquiries into uncertainty to increasingly formalised methods for planning empirical research.

This article traces these developments through key periods and thinkers, beginning with Christiaan Huygens quantitative framing of uncertainty and Jacob Bernoulli’s pioneering work on the ‘law of large numbers’. It then follows the maturation of frequentist sample size calculation through hypothesis testing and experimental design, alongside later Bayesian approaches to uncertainty and planning, while briefly noting areas in which methodological development remains ongoing.
The origins of probabilistic thinking
Christiaan Huygens was among the first to frame uncertainty in quantitative terms. In De Ratiociniis in Ludo Aleae (Huygens 1657), Huygens sought to understand how fair outcomes could be mathematically reasoned in games of chance. Building on earlier correspondence between Blaise Pascal and Pierre de Fermat in 1654, Huygens introduced the concept of expected value as a rational basis for decision-making under uncertainty. Although his work did not address sample size directly, Huygens’ ideas established a new way of thinking about randomness, thereby shifting probabilistic inquiry from philosophy to computation, and laying the groundwork for future efforts to relate empirical data to underlying truths.
Jacob Bernoulli and the ‘law of large numbers’
The roots of sample size estimation can be traced to Jacob Bernoulli, a Swiss mathematician from the famous Bernoulli family. In his posthumously published masterpiece, Ars Conjectandi (Bernoulli 1713), Bernoulli articulated one of the most important results in probability theory: the law of large numbers, laying the mathematical foundations for the view of probability as a long-run relative frequency over an infinite series of repeated trials.
Bernoulli was interested in how many observations were needed to estimate the true probability of a success in repeated, independent experiments. For example, how many times a coin must be flipped to ensure that the observed proportion of heads approximates the true probability. He formalised this question using ‘Bernoulli trials’ – experiments with binary outcomes (success or failure) – and showed that the observed proportion converges to the true probability as the number of trials increases.
Bernoulli went further by attempting to derive a numerical bound on the number of trials required to estimate an unknown probability with a specified degree of accuracy and confidence. In Ars Conjectandi (Bernoulli 1713), he posed what later became known as ‘Bernoulli’s problem’, applying his theory to the question at the core of sample size determination: how many observations are needed to ensure that the observed proportion of successes lies within a given margin of the true probability with high probability. Bernoulli illustrated this using an analogy of drawing coloured stones from a jar, where repeated random draws progressively reveal the underlying proportion, much like more patients entering a clinical trial allow estimation of the true treatment response proportion with increasing precision.
However, the numerical bounds Bernoulli derived for this purpose were both extremely conservative and methodologically crude. As Karl Pearson later demonstrated in a detailed historical and mathematical reassessment, Bernoulli relied on a loose system of inequalities rather than analytic approximations to the binomial distribution, resulting in sample size requirements that were exaggerated by several fold relative to later solutions (Pearson 1925). In the case of Bernoulli’s original jar example, to estimate a probability of 0.6 within ±0.02 with odds of 1000:1, Bernoulli’s method yielded a required sample size of more than 25,000 observations, a result so impractical that it seems to have deterred Bernoulli from publishing Ars Conjectandi; it appeared eight years after his death.
Importantly, the inverse square root relationship between estimation accuracy and sample size, frequently attributed to Bernoulli in later textbooks, does not appear in Ars Conjectandi. As Karl Pearson showed, this relationship emerges only with Abraham de Moivre’s normal approximation to the binomial distribution, developed in the 1730s and later formalised through the central limit theorem (de Moivre 1738; Pearson 1925). The subsequent attribution of this result to Bernoulli reflects a historical conflation rather than a faithful reading of the original work.
Bernoulli’s lasting contribution therefore lies not in providing a practical solution to sample size determination, but in clearly formulating the problem itself: the recognition that a quantitative relationship exists between sample size, uncertainty, and the reliability of empirical inference. This conceptual insight that large samples may be required to achieve high degrees of certainty remains a foundational lesson for statistical reasoning, even though the mathematical tools needed to operationalise it were supplied by later developments.
The 18th–19th century: probability theory matures, but sample size lags
After Bernoulli, the 18th and 19th centuries saw a flowering of probability theory, with major contributions from Pierre-Simon de Laplace, Carl Friedrich Gauss, and Siméon-Denis Poisson. Laplace developed early versions of the ‘central limit theorem’ and used normal approximations to estimate probabilities (Laplace 1812). Gauss introduced the method of least squares and assumed normally distributed errors in measurements (Chow et al. 2008). Poisson developed his namesake distribution to model rare events in large populations (Hanley 2022). Further, several prominent researchers, including Jules Gavarret and Thomas Graham Balfour saw the dangers later called type I and type II errors in small numbers (Campbell 2013).
Despite these advances, sample size estimation remained at best implicit, and even then, approximate and based on pragmatic considerations such as availability of patients or therapy. Most applications involved censuses or observational studies with fixed sample sizes. There was little incentive or infrastructure to ask how many observations were enough.
The Neyman–Pearson revolution and the formalisation of hypothesis testing (1920s–1930s)
A turning point came with the work of Jerzy Neyman and Egon Pearson in the 1920s and 1930s. Their framework introduced formal concepts of hypothesis testing, type I and type II errors, and statistical power (Neyman and Pearson 1933). These innovations enabled researchers to calculate the minimum sample size needed to achieve a desired probability of detecting a true effect, given an acceptable false-positive rate.
Their work culminated in the Neyman–Pearson ‘lemma’ (Neyman and Pearson 1933), which provided a general theoretical foundation for constructing the most powerful tests for simple hypotheses. Although the Neyman–Pearson framework itself is distribution-agnostic, its application to commonly used parametric settings led, in subsequent methodological work, to closed-form sample size equations for standard scenarios. For example, in comparing two population means under a normality assumption with unknown variance, the required sample size per group is often approximated as:

where δ is the minimum detectable difference between groups, σ is the assumed standard deviation, and σ is the type I error rate (probability of a false positive) and β is the type II error rate (probability of a false negative). Although this formula is derived using normal-theory approximations, it is routinely used to plan studies analysed using two-sample t-tests, particularly when σ must be estimated from the data.
Similarly, for tests of proportions, widely used sample size formulas are based on large-sample normal approximations to the binomial distribution (Cochran 1953; Cochran and Cox 1957), for example:

where p₁ and p₂ are the proportions in the two groups, and p is a pooled or average proportion (often taken as (p₁ + p₂)/2) used in the variance approximation under the null hypothesis. These formulas allowed researchers to pre-specify sample sizes based on expected effect sizes, substantially advancing the prospective planning of clinical trials, survey experiments, and industrial quality testing. Notably, these formulations make explicit the problem that Bernoulli had only glimpsed: that achieving higher levels of precision or confidence requires disproportionately larger sample sizes.
From a historical perspective, both formulas implicitly rely on the law of large numbers and the central limit theorem (unless normality is already assumed). The law of large numbers ensures that, as the sample size increases, the sample mean or proportion converges to its true population value, making the estimates more stable. The central limit theorem is used to justify approximating the sampling distribution of these statistics by a normal distribution, even when the underlying data are not perfectly normal. This normal approximation is what allows the use of z-scores in the formulas and makes closed-form sample size calculations feasible (direct calculation of required sample sizes using specific algebraic formulas).
The broader conceptual shift introduced by Neyman and Pearson was therefore not only technical but also methodological. Statistical inference was no longer treated as a purely post hoc exercise applied after data collection, but a prospective design choice requiring explicit specification of acceptable error rates, effect sizes, and study objectives. This reorientation made it possible to determine sample sizes in advance in a principled and reproducible way, aligning study design with scientific aims, ethical considerations, and practical constraints. Neyman-Pearson-based sample size calculations thus formalised and approach to study planning focused on controlling inferential error rates when testing clearly specified hypotheses.
Fisher’s influence and the rise of experimental design
In parallel with Neyman and Pearson, Sir Ronald Aylmer Fisher revolutionised empirical science by formalising the principles of experimental design in agriculture and biology. In his seminal work, Design of Experiments (Fisher 1935), Fisher emphasised techniques like randomisation, blocking, and factorial designs to reduce bias and improve precision. His introduction of the analysis of variance enabled researchers to partition variability and test multiple hypotheses simultaneously.
Although Fisher did not advocate for power analysis in the way that Neyman and Pearson did, his methods implicitly addressed data sufficiency. For example, in ANOVA, the F-statistic compares the variance between groups to the variance within groups:

where MS denotes the mean square. The non-centrality parameter of the F-distribution, which is a function of sample size and effect size, determines the power of the test. As later statisticians such as Oscar Kempthorne as well as William Gemmell Cochran and David Roxbee Cox expanded Fisher’s designs, they derived explicit sample size formulas to detect main effects or interactions in balanced factorial experiments (Kempthorne 1952; Cochran and Cox 1957).
For instance, in a one-way ANOVA with k groups and equal group sizes n, the sample size needed to detect a standardised effect f with significance level ɑ/2 and power 1-β for a two- sided test is approximately:

Fisher’s emphasis on efficiency, replication, and systematic control of variability laid the conceptual and mathematical groundwork for sample size optimisation. While Fisher favored inductive inference and resisted Neyman-Pearson’s binary decision framework, his contributions were foundational to later developments in power analysis, resource allocation, and design-based sample size planning.
Early references to sample size estimation for clinical studies
Early efforts to translate statistical principles into practical guidance for clinical research can be found in the work of Donald Mainland and Austin Bradford Hill in the 1930s (Mainland 1938; Hill 1937). Mainland had already addressed problems of chance and numerical adequacy in clinical work in a 1936 BMJ paper, prepared with statistical advice from Fisher, before expanding these ideas in his 1938 book (Mainland 1936; Mainland 1938). Although neither author presented explicit formula-based derivations, both provided numerical guidance on the relationship between sample size, detectable effect, and statistical uncertainty that is broadly consistent with modern theory.
In his 1938 book, Mainland considered both small studies, using exact probability arguments, and larger studies, where normal approximations were invoked 1930s (Mainland 1938. His calculations, developed with input from Fisher, illustrate an early attempt to quantify the number of observations required to detect differences of practical relevance under varying assumptions. Hill, in the first edition of Principles of Medical Statistics, focused on larger studies and provided approximate sample size requirements for detecting effects of different magnitudes (Hill 1937). His recommendations correspond closely to conventional significance thresholds and modest power, often effectively targeting a probability of detection near 50% and rounding sample sizes upward for practical use.
Taken together, these contributions represent an early ‘broad-brush’ approach to sample size planning, where the emphasis was on approximate detectability rather than formal error control or explicit optimisation. They also highlight that, despite the availability of theoretically grounded methods, sample size determination in clinical research remained largely pragmatic, shaped by feasibility and context rather than strict adherence to formal calculation.
The post-World War II era (1): operational research, clinical trials, and standardised procedures (1940s–1960s)
By the post-World War II period, statistical thinking had already begun to permeate areas such as public health and agriculture, particularly following Fisher’s work at Rothamsted, and was increasingly extending into clinical medicine. The rise of randomised trials marked a pivotal shift toward rigorous causal inference in medicine. Earlier controlled studies, including quasi-randomised or alternation-based designs, had been conducted in the nineteenth and early twentieth centuries, illustrating growing awareness of the need to reduce bias, although true randomisation was not yet consistently implemented (Hróbjartsson et al. 1998; Stolberg 2006). This shift is generally regarded as having begun with the 1948 Medical Research Council streptomycin trial whose principal statistician was Austin Bradford Hill (Medical Research Council, 1948). As noted above, Hill had pointed out the need to address sample size issues in his hugely influential textbook “Principles of Medical Statistics” (Hill 1937). Despite this, however, there is no explicit mention of the calculation of sample size in the 1948 trial, which – like subsequent randomised trials – appears to have been dictated primarily by pragmatic considerations. Even the landmark 1954 Salk poliomyelitis vaccine trial determined its sample size pragmatically on epidemiological and logistical grounds, without reference to formal power calculations or detectable-effect criteria (Francis et al. 1955). Consistent with this broader pattern, early quantitative reasoning about sample size in clinical trials most often appeared retrospectively, through post hoc power calculations or illustrative discussions of detectable differences, rather than as prospective, protocol-defining design criteria. Table 1 provides early illustrative examples of randomised trials.

By the 1950s, practical planning tools for clinical and observational research began to become more widely available. Cochran’s Sampling Techniques (Cochran 1953) provided explicit sample size formulas for estimating population means and proportions with specified precision. For example, the sample size n required to estimate a population mean with margin of error E and standard deviation s at confidence level 1−ɑ/2 is:

This formula became widely used in public health surveys and clinical studies, where estimation rather than hypothesis testing was the main goal. For binary outcomes, Cochran and others popularised the normal approximation to the binomial distribution for planning prevalence surveys:

where p is the anticipated proportion and E is the desired margin of error. These equations enabled pre-study planning based on anticipated variability, target precision, and confidence thresholds.
The Post-World War II era (2): ethical rigor, sequential efficiency, and the contributions of Hill and Wald (1940s–1960s)
In addition to advocating for the use of randomisation and blinding in clinical trials for the elimination of allocation bias, Hill emphasised the ethical and scientific imperative of pre-specifying trial objectives and statistical parameters, including sample size, to avoid misleading post hoc interpretations. He outlined practical guidance for calculating sample sizes based on detectable treatment differences and clinically meaningful effect sizes, long before these practices were codified in regulatory guidance. Hill’s influence extended globally, shaping the methodology of early therapeutic trials in cardiology and infectious diseases.
Meanwhile, in the realm of operational research, Abraham Wald developed the theory of sequential analysis, which provided a radically different approach to determining how much data to collect. Instead of fixing the sample size in advance, Wald’s method allowed data to be evaluated as it accumulated, with formal rules for stopping early if results were sufficiently strong. His most well-known contribution, the sequential probability ratio test (SPRT), minimised the average number of observations needed while controlling error rates (Wald 1947). For example, in the binary case, Wald derived boundary conditions for likelihood ratios such that sampling could be stopped as soon as the accumulated evidence crossed a critical threshold. Though originally applied in industrial quality control during World War II (Wald 1943), SPRT and its extensions were soon adapted to clinical trials, vaccine testing, and epidemiological surveillance, particularly in settings where minimising patient burden or cost was paramount (Armitage 1957). The work further set the stage for the group sequential design methodologies such as O’Brien-Fleming monitoring boundaries (O’Brien and Fleming, 1979) and Lan-DeMets (Lan and DeMets, 1983) alpha-spending that are now used so widespread in large well-designed clinical trials (Gluud and Thorlund 2025).
Together, Hill and Wald demonstrated two complementary philosophies of sample size planning: one grounded in fixed-sample ethical rigor and trial design, the other in adaptive efficiency and probabilistic control. Both helped to cement the idea that sample size is not merely a logistical detail but a central element of scientific validity. At the same time, these post-war developments also exposed a broader conceptual issue that would later be articulated more explicitly. While both fixed-sample and sequential methods refined control of statistical errors and efficiency, they largely retained an explanatory orientation, focusing on hypothesis testing and estimation under controlled conditions. In a seminal paper, Schwartz and Lellouch subsequently distinguished between explanatory trials, designed to test biological hypotheses, and pragmatic trials, designed to support treatment decisions under real-world conditions, a distinction with direct implications for how sample size should be determined (Schwartz and Lellouch, 1967).
Statistical standardisation, expansion by outcome types, and the era of software (1970s–1990s)
By the 1970s, the field of sample size estimation became increasingly standardised, driven by widely adopted textbooks and statistical training. The most influential work during this period was Jacob Cohen’s Statistical Power Analysis for the Behavioral Sciences (Cohen 1977; revised 1988), which formalised the use of effect size conventions and provided ready-to-use power tables for common designs. Cohen introduced thresholds for small (0.2), medium (0.5), and large (0.8) standardised mean differences (Cohen’s d), which could be paired with formulas such as:

to estimate the sample size required for a two-sample t-test under standard normal assumptions. He also provided analogous formulas for correlations, proportions, and ANOVA, enabling psychologists and social scientists to conduct a priori power analyses even without deep statistical training.
As sample size became embedded in applied research, general principles of precision and error control increasingly gave rise to outcome-specific formulas tailored to different types of outcomes. In settings characterised by varying follow-up times or recurrent events, attention shifted from risks to event rates, naturally leading to models based on counts of events observed over person-time. These settings are well described by the Poisson distribution, which provides a likelihood-based framework for incidence rates.
By the mid 1970s to early 1980s, methodological work had extended to models and closed form formulas for event counts and incidence rates, typically derived using large-sample approximations to the distribution of estimated rates or log rate ratios (Gail 1974; Brown and Green 1982). Related developments addressed time-to-event outcomes, where sample size and power calculations were formulated in terms of the expected number of events rather than the number of participants, drawing on asymptotic properties of logrank tests and proportional hazards models (Schoenfeld 1981; Schoenfeld 1983; Freedman 1982), which were increasingly cited in applied trial planning. Some examples of early clinical trials employing formal sample size planning and calculations for different outcome types are shown in Table 1. To our knowledge, it is not currently possible to identify with certainty the first clinical trial that employed a formal a priori sample size calculation, reflecting both incomplete reporting and the gradual adoption of these methods in practice.
Concurrently, regulatory agencies began to formalise the requirement that clinical trials justify their sample size based on expected effect size, variance, significance level, and desired power. Key milestones included the International Conference on Harmonization (ICH) E9 guideline (ICH 1996) and the Food and Drug Administration’s (FDA’s) Guidance for Industry documents (FDA 1998), which mandated transparent documentation of sample size logic in trial protocols. These documents shifted sample size estimation from an academic best practice to a regulatory obligation. This institutional shift was reinforced by later meta-epidemiological work showing that smaller randomised trials, on average, report lower methodological quality and yield systematically different treatment effect estimates than larger trials in meta-analyses (Kjaergard et al. 2001; Gluud et al. 2008). At the same time, evidence suggests that the adoption of formal sample size calculations often preceded their routine reporting in the medical literature (Soares et al. 2004).
The emergence of dedicated statistical software in the 1980s and 1990s played a transformative role in the dissemination and standardisation of sample size estimation (Table 2). Commercial platforms automated increasingly complex calculations, facilitated exploration of alternative design assumptions, and expanded access to advanced methods such as interim monitoring, sample size re-estimation, and adaptive trial designs. Collectively, these tools reduced technical barriers and consolidated sample size planning as a routine component of applied research across disciplines. As sample size estimation became embedded within regulatory, educational, and computational infrastructures, frequentist approaches reached a high degree of methodological maturity and routinisation.
The 2000s to the present: from computational flexibility to open-source accessibility
In the early 2000s, advances in computing power and statistical software made it practical to apply simulation-based methods for sample size estimation in increasingly complex study and trial designs. Monte Carlo simulation and resampling techniques were widely adopted to evaluate operating characteristics such as power and precision in hierarchical, nonlinear, and time-to-event models, particularly in settings where closed-form planning formulas were unavailable (Chow et al. 2008). Although initially developed largely within a frequentist framework, these computational approaches also provided the foundation for later Bayesian planning methods.
During this period, Bayesian approaches to sample size determination gained wider adoption, with planning criteria typically defined in terms of posterior informativeness rather than fixed error rates. Common strategies included targeting a desired posterior credible interval width or ensuring a high posterior probability that a treatment effect exceeded a clinically meaningful threshold (Gelman et al. 2013). These methods generally focused on inferential precision and evidence thresholds, rather than explicit optimisation of expected utility or loss.
In parallel, adaptive and group-sequential designs saw increasing use in the first decade of the twenty-first century, supported by increasingly sophisticated software platforms (Table 2). These developments broadened prospective study planning beyond fixed, closed-form formulas through interim analyses, sample size re-estimation, and other adaptive methodologies, while largely retaining an inferential rather than decision-analytic orientation.

In recent years, the rapid expansion of open-source statistical ecosystems, particularly R and Python, together with increasingly accessible graphical user interfaces, has transformed access to advanced sample size and power analyses (Table 2). Methods that were once largely confined to specialised commercial software, such as simulation-based estimation, Bayesian planning, adaptive designs, and trial simulation, have become widely available through community-developed packages and web-based applications. The net effect has been the democratisation of advanced trial planning methods, making rigorous sample size estimation accessible to a global audience without requiring costly licences or advanced programming expertise. This trend shows no sign of abating, with continued development of open-source tools likely to further expand accessibility and functionality.
Taken together, however, these developments suggest an uneven maturity across frameworks. Frequentist sample size methods for explanatory trials are highly standardised, and Bayesian approaches have increasingly mature tools for planning based on posterior precision or decision thresholds. By contrast, as of 2026, sample size determination for explicitly pragmatic, decision-analytic trials, where utilities, costs, and consequences must be weighed, remains comparatively underdeveloped, with few widely adopted methods or software implementations (Harari et al. 2018; Austin et al. 2021). Related Bayesian developments have also begun to incorporate uncertainty in model performance and decision-relevant criteria such as assurance probabilities and value of information, particularly in the context of prediction model validation (Sadatsafavi et al. 2026). Lastly, recent work has begun to address this gap by proposing decision-analytic frameworks that link sample size to acceptable regret and the probability of incorrect treatment selection, often yielding substantially smaller or more interpretable sample sizes in pragmatic settings (Hozo et al. 2026).
Concluding remarks
In today’s data-rich environment, the central question is no longer simply ‘how much data is enough’, but rather ‘enough for what purpose’? For explanatory research, the statistical foundations of sample size estimation are largely mature. For pragmatic, decision-oriented studies, however, the problem remains open, echoing Schwartz and Lellouch’s call to align design, analysis, and sample size with the ultimate goal of the investigation. Future methodological and reporting guidance (e.g. CONSORT) may benefit from more explicitly recognising these distinctions between explanatory and pragmatic objectives when discussing sample size determination.
From Huygens’ philosophical exploration of uncertainty, through Bernoulli’s formalisation of the law of large numbers, to the statistical revolutions of the 20th century, the history of sample size estimation reflects the broader evolution of probability and applied statistics. Each era brought conceptual and practical advances: Neyman and Pearson’s framework transformed hypothesis testing into a tool for pre-study planning; Fisher and his successors established principles of design efficiency; and post-war methodologists linked statistical rigor to real-world constraints in medicine, economics, and other disciplines. This institutionalisation was further reinforced by the introduction of reporting standards such as the CONSORT statement (Begg et al. 1996; Moher et al. 2001; Schulz et al. 2010), which require explicit reporting of sample size calculations in randomised trials. Despite these requirements, empirical evaluations suggest that a substantial proportion of published trials still lack complete or reliable sample size justifications, even in high-impact journals (Charles et al. 2009).
The 21st century has built on these foundations with more advanced approaches such as adaptive designs (e.g. early stopping or sample size re-estimation), Bayesian approaches including informative priors sample size evaluations, as well simulation-based approaches to address more complex sample size estimations. Advances in computing and open-source ecosystems have expanded access to powerful analytical tools, lowering barriers for researchers across disciplines.
Collectively, these developments have transformed sample size estimation into a multifaceted discipline that informs inference, prediction, and decision-making. Rather than converging on a single universal principle, modern sample size planning increasingly reflects trade-offs among statistical precision, feasibility, ethical constraints, and decision relevance. As methods continue to evolve, the enduring challenge is not merely technical optimisation, but thoughtful integration of statistical theory with the practical and ethical realities of scientific inquiry.
For more information on the importance of having an adequate sample size for a clinical trial, see this entry in the Catalogue of Bias: Catalogue of Bias Collaboration, Spencer EA, Brassey J, Mahtani K, Heneghan C (2017). Wrong sample size bias. LINK
Acknowledgements
We thank Sarah Klingenberg (Copenhagen Trial Unit) for her diligent support in identifying relevant ‘early’ randomised trials for Table 1. We also thank Ofir Harari (Redwood AI) for his thorough review of later versions of the manuscript, ensuring accuracy and correctness in the historical description of statistical and probabilistic theory. We also thank the two reviewers (Benjamin Djulbegovic and Robert Matthews) and the James Lind Library editors for their many helpful suggestions.
References
Armitage P (1957). Restricted sequential procedures. Biometrika 44(1/2), 9–26.
Austin PC, Sapp RJ, Tu JV (2021). Informing power and sample size calculations when using propensity-score weighting with survival outcomes. Statistics in Medicine 40(19): 4317–4332.
Begg C, Cho M, Eastwood S, Horton R, Moher D, Olkin I, Pitkin R, Rennie D, Schulz KF, Simel D, Stroup DF (1996). Improving the quality of reporting of randomized controlled trials: The CONSORT statement. JAMA 276(8): 637–639.
Bernoulli J (1713). Ars Conjectandi (The Art of Conjecturing). Basel: Thurneysen Brothers.
Brown CC, Green SB (1982). Additional power computations for designing comparative Poisson trials. American Journal of Epidemiology 115(5): 752-8.
Campbell MJ (2013). Doing clinical trials large enough to achieve adequate reductions in uncertainties about treatment effects. Journal of the Royal Society of Medicine 106:68-71.
Charles P, Giraudeau B, Dechartres A, Baron G, Ravaud P (2009). Reporting of sample size calculation in randomised controlled trials: Review. BMJ 338: b1732.
Chouinard G, Annable L, Turnier L, Holobow N, Szkrumelak N (1985). A double-blind randomized clinical trial of lithium carbonate and haloperidol in the treatment of acute mania.
Biological Psychiatry 20(4): 353–365.
Chow S-C, Shao J, Wang H (2008). Sample Size Calculations in Clinical Research (2nd ed.). Boca Raton, FL: Chapman and Hall/CRC.
Cochran WG (1953). Sampling Techniques. New York: John Wiley and Sons.
Cochran WG, Cox GM (1957). Experimental Designs (2nd ed.). New York: John Wiley and Sons.
Cohen J (1977). Statistical Power Analysis for the Behavioral Sciences (1st ed.). New York: Academic Press.
Cohen J (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Earl HM, Rudd RM, Spiro SG, Ash CM, Geddes DM, Souhami RL (1991). A randomised trial of etoposide and carboplatin versus etoposide and cisplatin in small-cell lung cancer.
British Journal of Cancer 63(5): 800–804.
FDA (1998). Guidance for Industry: E9 Statistical Principles for Clinical Trials. U.S. Department of Health and Human Services. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/e9-statistical-principles-clinical-trials
Fisher RA (1935). The Design of Experiments. Edinburgh: Oliver and Boyd.
Francis T Jr, Korns RF, Voight RB, Boisen M, Hemphill FM, Napier JA (1955). An evaluation of the 1954 poliomyelitis vaccine trials: Summary report. American Journal of Public Health 45(5 Pt 2): 1–63.
Freedman LS (1982). Tables of the number of patients required in clinical trials using the logrank test. Statistics in Medicine 1(2): 121-129.
Gail M. (1974). Power computations for designing comparative Poisson trials. Biometrics 30: 231–237.
Gelman A, Carlin JB, Stern HS, Dunson D B, Vehtari A, Rubin DB (2013). Bayesian Data Analysis (3rd ed.). Boca Raton, FL: CRC Press.
Gluud LL, Thorlund K, Gluud C, Woods L, Harris R, Sterne JAC (2008). Correction: Reported methodologic quality and discrepancies between large and small randomized trials in meta-analyses. Annals of Internal Medicine 149(3): 219.
Gluud C, Thorlund K (2025). Long overdue recognition of Klim McPherson’s 1974 article on sequential analysis of trial data. Journal of the Royal Society of Medicine 118:304-307.
Hanley JA, Bhatnagar S (2022). The “Poisson” distribution: History, reenactments, adaptations. The American Statistician 74(2): 363-371.
Harari O, Bingham DR, Dean A, Higdon D (2018). Computer experiments: Prediction accuracy, sample size and model complexity revisited. Statistica Sinica 28(2): 899–919.
Hill AB (1937). Principles of Medical Statistics. London: The Lancet.
Hozo I, Hemkens LG, Djulbegovic B. (2026). Sample size determination for decision-centered pragmatic trials. Journal of Clinical Epidemiology (published online ahead of print) doi: 10.1016/j.jclinepi.2026.112397.
Hrobjartsson A, Gøtzsche PC, Gluud C (1998). The controlled clinical trial turns 100 years: Fibiger’s trial of serum treatment of diphtheria. BMJ 317: 1243.
Huygens C (1657). De Ratiociniis in Ludo Aleae [On Reasoning in Games of Chance]. (English translation by Thomas O. Beebee (2006)). In F.N. David (Ed.), Games, Gods and Gambling (pp. 69–74). Dover Publications.
International Conference on Harmonisation (ICH) (1996). E9 Statistical Principles for Clinical Trials. (https://www.ema.europa.eu/en/documents/scientific-guideline/ich-e-9-statistical-principles-clinical-trials-step-5_en.pdf)
Kempthorne O (1952). The Design and Analysis of Experiments. New York: John Wiley and Sons.
Kjaergard LL, Villumsen J, Gluud C (2001). Reported methodologic quality and discrepancies between large and small randomized trials in meta-analyses. Annals of Internal Medicine 135(11): 982–989.
Lan KKG, DeMets DL (1983). Discrete sequential boundaries for clinical rials. Biometrika 70: 659-663.
Laplace PS Marguis de (1812). Théorie analytique des probabilités. Paris: Courcier.
Mainland D (1936). Problems of chance in clinical work. British Medical Journal 2(3943): 221–224.
Mainland D (1938). The Treatment of Clinical and Laboratory Data: An Introduction to Statistical Ideas and Methods for Medical and Dental Workers. Edinburgh: Oliver and Boyd.
Medical Research Council (1948). Streptomycin treatment of pulmonary tuberculosis: a Medical Research Council investigation. BMJ 2(4582): 769–772.
Moher D, Schulz KF, Altman DG (2001). The CONSORT statement: Revised recommendations for improving the quality of reports of parallel-group randomized trials. JAMA 285(15): 1987–1991.
de Moivre A (1738). The Doctrine of Chances: Or, A Method of Calculating the Probabilities of Events in Play. 2nd ed. London: W. Pearson.
Neyman J, Pearson ES (1933). On the problem of the most efficient tests of statistical hypotheses. Philosophical Transactions of the Royal Society A 231(694–706): 289–337.
O’Brien PC, Fleming TR (1979). A multiple testing procedure for clinical trials. Biometrics 35: 549-556.
Pearson K (1925). James Bernoulli’s Theorem. Biometrika 17(3/4): 201-210.
Sadatsafavi M, Gustafson P, Setayeshgar S, Wynants L, Riley RD (2026). Bayesian sample size calculations for external validation studies of risk prediction models. Statistics in Medicine 45: e70389.
Scandinavian Simvastatin Survival Study Group (1994). Randomised trial of cholesterol lowering in 4,444 patients with coronary heart disease: The Scandinavian Simvastatin Survival Study (4S). The Lancet 344(8934): 1383–1389.
Schwartz D, Lellouch J (1967). Explanatory and pragmatic attitudes in therapeutic trials. Journal of Chronic Diseases 20: 637-648.
Schoenfeld D (1981). The asymptotic properties of nonparametric tests for comparing survival distributions. Biometrika 68(1): 316–319.
Schoenfeld DA (1983). Sample-size formula for the proportional-hazards regression model. Biometrics 39: 499–503.
Schulz KF, Altman DG, Moher D (2010). CONSORT 2010 statement: Updated guidelines for reporting parallel group randomized trials. BMJ 340: c332.
Soares HP, Daniels S, Kumar A, Clarke M, Scott C, Swann S, Djulbegovic B (2004). Bad reporting does not mean bad methods for randomised trials: observational study of randomised controlled trials performed by the Radiation Therapy Oncology Group. BMJ 328(7430): 22–24.
Stolberg M (2006). Inventing the randomized double-blind trial: The Nürnberg salt test of 1835. Journal of the Royal Society of Medicine 99:642-643.
Wald A (1943). Sequential tests of statistical hypotheses. Annals of Mathematical Statistics 14(2): 117–186.
Wald A (1947). Sequential Analysis. New York: John Wiley & Sons.
Werzberger A, Mensch B, Kuter B, et al. (1992). A controlled trial of a formalin-inactivated hepatitis A vaccine in healthy children. New England Journal of Medicine 327(7): 453–457.
Zubrod CG, Schneiderman M, Frei E III, et al. (1960). Appraisal of methods for the study of chemotherapy of cancer in man: Comparative therapeutic trial of nitrogen mustard and triethylene thiophosphoramide. Journal of Chronic Diseases 11(1): 7–33.
