Skip to content
AI.info

Causal inference

Power, Minimum Detectable Effects, and Sample Size

Connect effect size, variance, alpha, power, clustering, repeated measures, attrition, and decision value in experiment planning.

By the end you can

Example

Sixty-seven of seventy-one "negative" trials had never tested anything

Seventy-one published randomised control trials had reported their treatment as no different from control. In 1978 Freiman and three colleagues went back through all of them, in the New England Journal of Medicine, and asked a question none of the original reports had asked. Given the number of patients actually enrolled, what were the chances that trial would have detected a real improvement if one had been there?

The answer: 67 of the 71 carried a greater than 10 per cent risk of having missed a true 25 per cent therapeutic improvement. In 57 of them the 90 per cent confidence interval was still compatible with a 25 per cent improvement. The published data would have looked much the same whether the treatment did nothing at all or raised the response rate by a quarter. The negative result was never in the data. It was a report on the sample size, read as a verdict on the therapy. The abstract says it plainly: “Many of the therapies labeled as "no different from control" in trials using inadequate samples have not received a fair test.”

Not one of those 71 outcomes was decided at the analysis stage. Four quantities settled what each trial could ever say. All four were fixed while the protocol was being written, whether or not anyone wrote them down.

  • Alpha is the false-alarm rate the team agrees to live with in advance — how often they are willing to call an effect real when nothing is there. ICH E9 puts the convention at 5% or less.
  • Power is the chance that this particular design would catch a particular effect, supposing that effect is real. Its complement is the type II error, which ICH E9 describes as “conventionally set at 10% to 20%”. Against a 25 per cent improvement, 67 of Freiman's 71 trials failed that standard.
  • The minimum detectable effect is the smallest result the study could reliably pick up; anything below it stays invisible no matter what is true.
  • Precision is how tightly the estimate is pinned down — the width of the interval the team ended up staring at. In 57 of those 71 trials that width still reached all the way to a 25 per cent improvement.

Power belongs to a design and one effect, not to an experiment

Power is not something an experiment has. You can compute it only after naming an effect size and fixing a design. The design is the whole list: which quantity you are estimating, how much the outcome varies, how units are split between arms, how many there are, whether they arrive in clusters or as repeated observations from the same source, how many are lost along the way, and what the analysis will do with whatever is left. Change any one of those and the number changes with it.

So the useful question is turned around. Decide first what the smallest effect worth acting on is — the point below which the decision owner would do nothing differently. Then ask whether the design can detect something that small. That single comparison is what a power analysis is for: the minimum detectable effect, set against the smallest effect that would change a decision. A study that can only see effects larger than the ones that matter does not answer the decision, however clean its output looks.

Whole research literatures skip the comparison. Every neuroscience meta-analysis published in 2011 — 49 of them, containing 730 individual studies — was gone through by Button and colleagues in 2013. For each study they computed its power to detect the effect its own meta-analysis had estimated. The median came out at 21%. Their paper in Nature Reviews Neuroscience opens with the finding itself: “Here, we show that the average statistical power of studies in the neurosciences is very low.” Nord and colleagues restated and reanalysed the figure in the Journal of Neuroscience in 2017, and it held. A typical design in that literature would miss the effect it was hunting roughly four times in five. Its authors did not know that when they ran it.

Astronomers make the comparison before choosing an instrument. The faintest object that matters is named first, and the telescope is picked to resolve it. One that cannot will still return beautifully sharp images of brighter objects, and still fail the mission. In an experiment the faint object's brightness is itself a guess. Effect sizes and variances are assumptions going in, not fixed properties of the sky. That is why a plan is written across a range of them rather than a single value.

None of the 71 trials Freiman's team examined had made the comparison either. Had they made it, they would have known before enrolling a patient that a 25 per cent improvement could come and go without leaving a mark on the result.

A result that clears no threshold is ambiguous by construction — small effect, noisy outcome, weak design, too few units — and it never says which.

Example

Row counts flatter the evidence you actually have

The reflex when a study looks thin is to collect more data. That works only when the extra rows carry independent information. Several ordinary features of a design quietly guarantee they do not.

Group people together and their outcomes stop being independent. How much? Adams and colleagues reanalysed 31 cluster-based primary care studies in 2004 and estimated intraclass correlation coefficients for 1,039 variables. Their Results read: “ICCs were estimated for 1,039 variables. The median ICC was 0.010 (interquartile range [IQR] 0 to 0.032, range 0 to 0.840).” The median looks negligible. The range does not. The same measure moved between datasets too: SF-36 physical functioning ranged 0.001–0.055 across six datasets, systolic blood pressure 0–0.052 across four. Their own conclusion is that the ICC “can rarely be estimated in advance”. So plan across a range, not a point.

Attrition is the other leak, and the standard repair does not work. A panel convened by the US National Research Council at the FDA's request reported on missing data in clinical trials in 2010. The habit it examined is inflating the sample to cover an expected dropout rate. The panel wrote: “This approach is generally flawed, since inflating the sample size accounts for a reduction in precision of the study from missing data but does not account for bias that results when the missing data differ in substantive ways from the observed data.” It is equally blunt about rescuing the study afterwards: “There is no 'foolproof' way to analyze data subject to substantial amounts of missing data”. Extra enrolment buys back precision. Nothing buys back randomisation.

What power depends on is not how many rows the table holds. It is how many independent pieces of evidence those rows amount to. The gap between the two can be enormous, and the row count never shows it.

  • When people are grouped — pupils inside schools, patients inside clinics — their outcomes move together. Take the median ICC of 0.010 that Adams and colleagues found, with 50 patients per practice. The design effect is 1 + (50−1) × 0.010 ≈ 1.49, so about 1,000 patients carry the information of about 670. At the upper end of their observed range, 0.840, a whole clinic is worth barely more than one patient.
  • Many events logged from the same user are one user seen repeatedly rather than many users, and the event count flatters the sample badly.
  • A rare outcome can leave a very large study with only a handful of the events that carry any signal. What matters is how often the thing happened, not how many rows were collected.
  • Whoever fails to come back for follow-up takes their information with them. Padding the enrolment to compensate is what the NRC panel called “generally flawed”: extra units buy back precision, and nothing buys back the unbiasedness randomisation was supposed to deliver.

Key idea

Observed power is the p value wearing different units

Once an interval comes back wide, someone will propose calculating the power the study "had", using the effect that was actually observed. That calculation carries no information at all. Hoenig and Heisey showed in 2001 that observed power is determined completely by the p value. A nonsignificant p value always corresponds to a low observed power, and in the Z-test case observed power comes out at exactly .5 when p equals alpha. Their section on observed power states the consequence plainly: “Computing the observed power after observing the p value should cause nothing to change about our interpretation of the p value.”

They found methodological recommendations advocating post-experiment power calculation across 19 applied journals, and at least two journals ask for such calculations as a matter of policy. The practice was not a private error. It was being required. And it is worse than empty, because it can dress a badly uncertain estimate up as a firm conclusion about the design, in either direction.

What a design review actually needs is the material the plan was built from: the interval around the estimate, the assumptions that fed the sample-size calculation, the range of effects the study was built to detect, and those calculations as they stood before any data existed. Every one of those can be argued with. Observed power cannot. It carries no content beyond the result it was derived from.

The design question was answered before the data existed, so only the numbers computed back then are in a position to answer it.

Example

Four words that get used as if they meant the same thing

Design reviews swap these four for one another routinely, and that swap is exactly what lets a study be signed off without anyone asking whether it can detect an effect worth acting on. Each of them answers a different question.

  • Power answers a conditional question: given this design, and an effect of this particular size, how often would the study detect it? The answer moves as soon as either condition moves. The 21% median Button and colleagues computed is that question asked 730 times, with each study's own meta-analytic effect supplied as the condition.
  • The minimum detectable effect turns that around. It reports the smallest effect the design could reliably catch, once the false-alarm rate and the power target have been chosen — 5% or less and 10% to 20% type II error, in the ICH E9 conventions.
  • The design effect is the price of structure. It is the factor 1 + (m − 1) × ICC by which the variance inflates when units arrive in clusters of size m instead of independently. At the median ICC of 0.010 measured across 1,039 primary care variables, clusters of 50 cost a factor of about 1.49 before the study has done anything wrong.
  • Practical significance has nothing to do with the other three. It is the size at which someone would genuinely act differently — the 25 per cent improvement in Freiman's survey, say. It comes out of the decision, not out of the data.

Steps

Only a third of trials in the best journals get this plan right

The profession is not good at this, and someone has counted. Two hundred and fifteen two-arm superiority randomised controlled trials appeared in six high-impact general medical journals in 2005–2006. Charles and colleagues went through every one and reported the tally in the BMJ in 2009. Ten trials (5%) reported no sample size calculation at all. Ninety-two (43%) left out parameters needed to check one. Of the 157 reports with enough detail to recompute, 47 (30%) came out more than 10% away from the sample size the trial had reported. The value assumed for the control group was off by more than 30% in 31% of articles. Their summary: “Only 73 trials (34%) reported all data required to calculate the sample size, had an accurate calculation, and used accurate assumptions for the control group.”

A prospective plan is built in the order the decision is made, not the order the software asks for inputs. Start with the decision owner. What would they do differently, and how large would the effect have to be before they did it? That figure, not a conventional power target, is what the design has to clear.

From there it is arithmetic and honesty about the design. Take the effect worth acting on, the variance you expect in the outcome, the way units will be assigned and grouped, and the share you expect to lose to attrition. Work out what sample size gives an acceptable chance of detecting that effect. Then check what interval width the same sample buys, because a design can be powered to detect an effect and still return an estimate too vague to act on. ICH E9 requires the method, every input estimate and the basis for it to be written into the protocol in advance. It also requires the sensitivity of the plan to those estimates to be shown as a range of sample sizes rather than a single number. That is the discipline a control-group assumption wrong by more than 30% in 31% of those articles was there to enforce.

Where the design outruns the formulas — clustered assignment with an ICC that “can rarely be estimated in advance”, repeated measures, an analysis plan with several stages — simulate it instead. Generate data under your assumptions. Run the analysis you actually intend to run, and count how often it recovers the effect you planted. That is the same question the formula answers, asked in a form that survives the complications.

FigureProcess · 5 steps
  1. 1

    Define the estimand

    Unit, outcome, horizon, and effect scale.

  2. 2

    Set decision threshold

    Smallest effect worth implementing.

  3. 3

    Estimate variance

    Use historical data with design-consistent units.

  4. 4

    Model design losses

    Clustering, attrition, noncompliance, and multiplicity.

  5. 5

    Simulate operating characteristics

    Power, interval width, false positives, and stopping.

Visual

Where a study's information leaks away

The budget has five inputs, and each of them can be spent badly. The decision threshold fixes what counts as an effect worth finding at all — a 25 per cent therapeutic improvement, in the trials Freiman's team surveyed. The outcome's own behavior sets what a single observation is worth: how often it occurs, and how much it varies. The assignment design decides whether those observations stand alone or move together in clusters, at an ICC that can run anywhere from 0 to 0.840. Data loss removes some of them after the fact, and not always at random. That is the leak the NRC panel showed cannot be plugged by enrolling more people. The analysis plan then determines how much of what survived is genuinely used: the same recorded outcomes yielded twice the effective sample once CUPED regressed out a pre-experiment covariate.

Trace the path from one stage to the next and the leak becomes locatable. You can see which stage drained the information before it ever reached the estimate.

FigureHierarchy · 5 levels
  • Decision threshold

    Smallest effect that changes action.

    • Outcome behavior

      Baseline rate, variance, skew, and autocorrelation.

      • Assignment design

        Ratio, clustering, blocking, and repeated exposure.

        • Data loss

          Attrition, missing outcomes, and noncompliance.

          • Analysis plan

            Estimator, covariate adjustment, multiplicity, and stopping.

Comparison

A sample-size calculation protects one goal at a time

Every sample-size calculation serves an objective, and the plan should say out loud which one. Detecting a specified effect is the first: size the study so that if the effect is at least that large, the study will usually find it. Estimating with a stated precision is a different requirement. It asks for an interval no wider than some width, whether or not that interval happens to exclude zero. Ruling out an effect large enough to matter is the third, and the one most often skipped. Yet a study that can honestly report the effect is smaller than the threshold for action has answered the decision, even having found nothing.

Regulators codified the distinction long ago. ICH E9 says: “The sample size of an equivalence trial or a non-inferiority trial (see Section 3.3.2) should normally be based on the objective of obtaining a confidence interval for the treatment difference that shows that the treatments differ at most by a clinically acceptable difference.” Not a significance test. An interval, and a margin someone had to justify in advance.

The three objectives regularly imply different sample sizes. Declaring which one you bought is what prevents the reading Freiman's team found 71 times over: designs that had never been sized to rule out a 25 per cent improvement, written up as though they had ruled one out.

FigureComparison · 3 columns

Detection power

Chance to reject a null for a specified effect.

  • Hypothesis-test framing
  • Needs alpha and effect
  • Can ignore decision value

Precision target

Desired confidence-interval width.

  • Estimate-centered
  • Works without point null
  • Direct uncertainty goal

Decision value

Expected benefit of information relative to cost.

  • Policy-centered
  • Includes action threshold
  • May favor no experiment

When an experiment is worth running

Run it when the sample you can realistically obtain will separate effects worth acting on from noise, within uncertainty and risk you can accept. That is the entire test. It is a feasibility question about a decision rather than a statistical exercise.

When the answer is no, the options are wider than giving up. Run longer. Measure the outcome better, or choose one that varies less or occurs more often. Cut variance through the design. Pool the evidence with what already exists. Or conclude that on this question, at this scale, an experiment is not currently the thing that will inform the decision — which is a finding in itself, and a far cheaper one than the alternative.

Cutting variance is the option that sounds like an accounting trick and is not. Deng and colleagues presented CUPED in 2013: take data collected before the experiment started and use it as a covariate. They ran it on Bing's experimentation platform, and the abstract reports the payoff: “The results on Bing's experimentation system are very successful: we can reduce variance by about 50%, effectively achieving the same statistical power with only half of the users, or half the duration.” The users were the same users. The metric was the same metric. Only the analysis changed, and the information budget doubled.

The alternative is the route those 71 trials took. Promise a definitive answer from a design whose information budget was never going to support one. Then treat the wide interval it produced — an interval that in 57 cases still reached a 25 per cent improvement — as though it had settled the matter.

The plan's real job is to say in advance whether this study can answer the question, clearly enough that nobody runs it when it cannot.

Key takeaways