Causal inference
Multiplicity, Subgroups, and Sequential Monitoring
Plan primary outcomes, subgroup hypotheses, familywise or false-discovery control, interim monitoring, and decision rules.
By the end you can
- Explain how multiple outcomes and repeated looks increase false-positive risk
- Distinguish confirmatory and exploratory subgroup analyses
- Use interactions rather than separate significance tests for heterogeneity
- Choose familywise, false-discovery, or hierarchical error control
Example
The win that eighty chances were always going to produce
A product team ran an experiment and watched twenty outcomes across four user segments. One morning the dashboard refreshed and one metric, in one segment, had crossed p < 0.05. They wrote it up as a targeted win: the feature works, for these users. The roadmap turned toward that segment.
Nobody had decided in advance which of those comparisons the experiment was about. Twenty metrics in four segments is eighty chances for something to look striking. With eighty chances, a result of that size is roughly what you should expect when the feature does nothing at all. The team had not found an effect. They had found the largest of eighty noisy readings and given it a name.
The waste was not only the work that followed. The same screen carried segments where the metric had moved the wrong way, and nobody wrote those up. The report had been assembled to explain the win.
That team can always insist the effect was real, because nobody knows what was true underneath. So the rest of this lesson works with cases where the answer was known in advance: a trial split by astrological birth sign, a stroke therapy that was literally a dice roll, one data set handed to 29 teams at once, and simulations run on data built to contain no effect at all.
- The primary outcome is the one confirmatory endpoint the main decision was tied to, and it has to be chosen before the data arrive.
- A multiplicity family is the set of claims whose errors you have agreed to control together.
- An interaction is a direct test of whether the effect really differs between groups, rather than two separate verdicts read side by side.
- An interim look is any analysis run before the planned information horizon, including the one you took because the dashboard happened to be open.
The eighty chances were built before any data arrived
None of those comparisons came from anywhere exotic. Every one was an ordinary decision someone made while setting the experiment up: which outcomes to track, how many arms to run, where to cut the segments, which covariates and exclusions to fit, how often to look before the end. Each choice multiplies the chances that something compelling appears.
So the question error control has to answer is not how many p-values are on the page. It is how many decisions those numbers were allowed to influence. Correct over the family that matches the decisions, not over whatever list happened to get printed.
It is tempting to believe this only bites when the underlying effect is weak. It does not. ISIS-2 randomised 17,187 patients in 417 hospitals. Aspirin cut five-week vascular mortality from 1016 of 8600 patients (11.8%) to 804 of 8587 (9.4%), 2p<0.00001. The Lancet published that in 1988. The investigators then split their own trial by the patients' astrological birth sign, cuts chosen precisely because they could not possibly matter. Peter Sleight, one of them, reported what came out: “Even in a highly positive trial such as ISIS-2 [3], in which the overall statistical benefit for aspirin over placebo was extreme (P <0.00001), division into only 12 subgroups threw up two (Gemini and Libra) for which aspirin had a nonsignificantly adverse effect (9% ± 13%)”. Twelve slices of one of the most certain results in cardiology made the drug look harmful in two of them. The slicing variable was a birth sign.
Subgroups therefore need more than error control. A claim that the effect differs between two groups is a claim about a difference, and it has to be tested as one. That means an interaction or heterogeneity analysis, run with enough information to detect the difference, and confirmed by replication before anyone builds on it. Significant here and not significant there is two separate readings. It is not evidence that the two differ. Gemini and Libra are what that pattern looks like when the effect is identical everywhere.
The way out is to write down, before the data, what each analysis is for. One or a small set of prespecified decisions carry the primary claim. Key secondary claims support the interpretation: the mechanism, the accompanying benefit. Safety and guardrail measures are watched with deliberately asymmetric priorities, because you want a low bar for noticing harm. Exploratory analyses are allowed and useful, and what they produce is a hypothesis for someone to confirm later. Interim decisions — stop, declare futility, ramp, adapt — get their own rules, fixed before the first look.
Multiplicity is a property of the process that produced the table, and the table itself cannot show it to you.
Analogy
A loud spike is what a wide scan produces
A receiver sweeps across a huge span of frequencies through a night of static. Somewhere in that sweep there is always a spike louder than the rest, and the wider the sweep and the longer the night, the louder that spike will be. Its height tells you almost nothing until you know how much was searched to find it.
Experiments differ in one way worth naming as you go. Their outcomes are related to each other and they are not equally important, so eighty comparisons are neither eighty independent channels nor eighty equal ones. The pull is the same either way. The strongest signal in a wide search is mostly evidence about the width of the search.
Ask how much was scanned before you ask how loud the spike was.
Comparison
Confirmation, discovery and ordered testing need different machinery
There is no single correction to apply, because these are not one job. If a claim is going to be acted on and a false one is expensive, control the chance of making any false claim at all across the family. You will miss some real effects in exchange. If the point of the analysis is to generate candidates worth studying next, control instead the proportion of your announcements that turn out wrong. That tolerates some errors and finds more of what is there. And if the claims are ordered — a primary decision first, then secondary ones that only matter if the primary held — test them in that order and let each one gate the next.
The method follows from the decision you are about to make. It does not follow from the number of rows in the output.
Regulators reason the same way, and they have put it in writing. The FDA finalised its guidance on multiple endpoints in clinical trials on 21 October 2022, five years after the draft it replaced. It refuses to let the family be read off the endpoint list: “Both the issues and methods that apply to multiple endpoints also apply to other sources of multiplicity, including multiple doses, time points, or study population subgroups.” Doses, time points and subgroups are the dashboard's segments and windows under other names.
FWER control
Limits probability of any false rejection.
- Strong confirmation
- Can be conservative
- Useful for few critical claims
FDR control
Limits expected false proportion among discoveries.
- Supports exploration
- Allows some false findings
- Needs declared family
Hierarchical testing
Tests claims in a prespecified order.
- Matches decision tree
- Preserves power
- Invalid after reordering
Example
Four terms that are not interchangeable, however alike they sound
Two of these control errors, one tests a difference, one budgets error across time. Reports blur them constantly, and almost always in the direction that makes the claim sound stronger than it is. The first two are not folklore either. The split between them was made in a single paper, in 1995.
- FWER is the probability that a family contains at least one false rejection, which is the strict standard and the right one when a single wrong claim does real damage.
- FDR is the expected share of the hypotheses you rejected that are false. It accepts some wrong answers in return for finding more of the true ones. Benjamini and Hochberg defined it in 1995 and proved a procedure that controls it. The paper has been cited more than 108,000 times: Semantic Scholar counts 108,722, OpenAlex 110,694. Their own summary of the change is one line: “It calls for controlling the expected proportion of falsely rejected hypotheses — the false discovery rate.”
- An interaction test asks directly whether the treatment effect differs across groups, and it is the only thing that can support a subgroup claim.
- Alpha spending divides the type-I error budget across the interim analyses, so that looking early costs something instead of being free.
Steps
Write the plan while you still cannot see the results
All of this is done before the data accumulate, because afterwards you cannot separate your own reasoning from what you have already seen.
Name the primary decision and the outcome tied to it. List the whole family of claims that decision belongs to, including the segment cuts and time windows you intend to look at. Mark which analyses are confirmatory and which are exploratory, and then hold to it. Fix the looks: when you will look, what the rule is at each look, and how much of the error budget that look spends. Connect every analysis on the list to a decision and a claim status, so that nothing arrives later without a label.
None of that is a counsel of perfection invented for this lesson. It has been a regulatory requirement, with published wording, since 5 February 1998. That is the date on ICH E9, the guideline that sets the default label for subgroups: “In most cases, however, subgroup or interaction analyses are exploratory and should be clearly identified as such; they should explore the uniformity of any treatment effects found overall.” Exploratory unless prespecified as confirmatory. The burden runs the opposite way from the one most reports assume.
E9 also closes the loophole in the word look. It defines the term so broadly that no glance escapes it: “An interim analysis is any analysis intended to compare treatment arms with respect to efficacy or safety at any time prior to formal completion of a trial.” Any analysis, at any time. Including the one taken because the dashboard happened to be open.
A plan like this takes an afternoon. It is the only thing that lets you point at a result afterwards and say it was not selected.
- 1
Define claim families
Primary, secondary, safety, subgroup, and exploratory.
- 2
Choose control
FWER, FDR, hierarchical, or descriptive reporting.
- 3
Specify subgroup tests
Interaction, minimum information, and replication.
- 4
Plan interim looks
Information times, alpha spending, futility, and safety.
- 5
Log deviations
Record additions, boundary changes, and unplanned analyses.
Key idea
No correction can undo a family chosen after the fact
The dangerous case is not the team that forgets about multiplicity. It is the team that applies a correction to a family they assembled after seeing the numbers.
The controlled demonstration of that was run on dice, and the BMJ published it on 24 December 1994. One line of the abstract describes the whole intervention: “44 randomised controlled trials of DICE therapy for stroke were performed (simulated by rolling different coloured dice; two trials per investigator).” No treatment effect could exist, because the treatment was a die. Pooled, the 44 trials gave an 11% (SD 11) reduction in the odds of death — the nothing you would expect. Then Counsell and colleagues did what a determined analyst does with a disappointing pooled result. They excluded the red-dice trials, and the trials of poor methodological quality. The same data now showed a 22% (SD 13) reduction, 2P = 0.09. Restricted further to the experienced trialists, it showed 39% (SD 17), 2P = 0.02.
Every step of that arithmetic was correct. The family it was computed over was smaller and tidier than the search actually run, and the therapy was a dice roll. Decide which comparisons count once the results are in. Move a segment boundary because the effect looks cleaner on one side of it. Quietly leave out the outcome that went the wrong way. Whatever formula you then apply is describing a different experiment.
The defences are unglamorous. Prespecify the primary decisions. Log every analysis anyone runs, including the ones that went nowhere. Keep exploration and confirmation in separate parts of the report, in separate language. And treat any important claim that groups differ as unfinished until someone has replicated it.
A correction is only as honest as the record of what was tried.
Example
The family is always larger than the dashboard
Counting the metrics on the screen gives you the start of the family, not the family. The team in the opening case counted twenty outcomes and four segments and got to eighty. Even that was generous to themselves, because it left out everything below.
Three psychologists priced these choices instead of deploring them. Simmons and colleagues ran 15,000 simulations per scenario on data containing no real effect, and measured what each ordinary analytic freedom does to a nominal 5% false-positive rate. Psychological Science published the result in 2011. Simply having two dependent variables to choose between takes that rate to 9.5%. All four of the freedoms they tested, used together, take it to 60.7%. Their own summary: “A researcher is more likely than not to falsely detect a significant effect by just using these four common researcher degrees of freedom.”
- Every way of slicing the users is a comparison: geography, platform, tenure, risk band, demographics, and whichever combinations of them someone tried. In the same simulations, the freedom to control for gender, or for its interaction with the condition, was on its own worth a false-positive rate of 11.7% where 5% was nominal.
- Every time window is another one, from day one to week one to month one, plus each repeated look at the dashboard in between. One extra look — collecting 10 more observations per cell after an initial peek — moved the same nominal 5% to 7.7%.
- Every specification counts too: a different covariate set, a different exclusion rule, a different estimator, each producing its own number out of the same data. In 2018 one soccer-referee data set and one question went to 29 independent teams, 61 analysts in all. Silberzahn and colleagues reported what came back: “Overall, the 29 different analyses used 21 unique combinations of covariates.” The odds ratios ran from 0.89 to 2.93 around a median of 1.31. And 20 teams (69%) found a significant positive effect while 9 (31%) did not. Identical data.
- So does every version, whether that means extra treatment arms or successive launches that reuse related data and quietly re-test the same idea. The freedom to drop one of three conditions after the fact was measured at 12.6%, again against a nominal 5%.
Report the search, not just the winner
Lead with the decision you committed to and with how uncertain it still is. That is the one part of the report that was not selected, and it is what a reader is entitled to weigh first.
Everything after it carries its own status. For each secondary and subgroup result, say which family it belongs to, whether the difference was tested as a difference, and how much information there really was in that slice. Exploratory findings get exploratory verbs. They suggest something and they motivate a next test. They do not show or confirm.
How often does that happen? A 2012 review in the BMJ counted. Sun and colleagues went through the randomised controlled trials published in 2007 in the core clinical journals defined by the National Library of Medicine, a representative sample in the authors' description. Of 207 trials reporting subgroup analyses, 64 (31%) made a subgroup claim about the primary outcome. Only 6 of those 64 claims (9%) had a statistically significant test of interaction behind them. Only 26 (41%) had clearly prespecified the hypothesis. And 54 of the 64 (84%) met four or fewer of the review's 10 credibility criteria. Their conclusion is flat: “Authors often claim subgroup effects in their trial report. However, the credibility of subgroup effects, even when claims are strong, is usually low.” When a paper says the effect was concentrated in a subgroup, the base rate says nine times in ten the difference was never tested as a difference.
Safety runs on the opposite logic. A harm signal that has not crossed a multiplicity threshold is still a harm signal. The asymmetry is deliberate, because the cost of acting on a false alarm is not the cost of missing a real one.
The team in the opening case had none of this. Had the primary decision been written down first, the same eighty comparisons would still have existed. The win would simply have arrived as what it actually was: one exploratory result, worth a second experiment.
A report that hides its search is asking to be trusted about the one thing it never recorded.
Key takeaways
- Every extra outcome, segment, specification and look is another chance for noise to arrive looking like a result: two dependent variables took a nominal 5% false-positive rate to 9.5%, and four such freedoms together to 60.7%.
- The family you control errors over should match the decisions you actually made and the searching you actually did. The FDA's 2022 guidance extends it past endpoints to doses, time points and study population subgroups.
- A claim that two groups differ has to be tested as a difference, not read off two separate verdicts: 12 birth-sign cuts of ISIS-2, a trial with 2p<0.00001 overall, produced two subgroups in which aspirin looked harmful.
- FWER, FDR and hierarchical testing answer different questions, and choosing between them is a decision about consequences. Benjamini and Hochberg separated FDR from familywise error in 1995.
- Sequential monitoring only works when the looks and the stopping rules are fixed before the data arrive. ICH E9 has counted any pre-completion comparison of arms as an interim analysis since 5 February 1998.
- Exploratory findings should be labelled as exploratory and replicated before anyone acts on them; in 207 trials reporting subgroup analyses, only 6 of the 64 primary-outcome subgroup claims (9%) rested on a significant interaction test.