Skip to content
AI.info

Causal inference

Sensitivity Analysis, Negative Controls, and Falsification

Use quantitative sensitivity, negative controls, placebo tests, and specification stress tests without treating them as proof of validity.

By the end you can

Example

The vaccine worked best in the weeks before it could work at all

A cohort of 72,527 people aged 65 and over was followed for 8 years to see what influenza vaccination did to their risk of death. During influenza season, the relative risk of death for vaccinated versus unvaccinated persons was 0.56 (95% CI 0.52-0.61). That is the kind of number programmes are built on.

Then the same analysis was run in the period before influenza season began — a window in which the vaccine could not yet have worked. The relative risk there was 0.39 (95% CI 0.33-0.47). After the season it was 0.74 (0.67-0.80). The apparent protection was strongest in the stretch of calendar where there was nothing to protect against, and it faded as the season passed.

The standard remedy made things worse. Adjusting for diagnosis-code covariates moved every estimate further from the null, not closer to it. Jackson and colleagues were blunt about that in 2006: “In this study, the magnitude of the bias demonstrated by the associations before the influenza season was sufficient to account entirely for the associations observed during influenza season.”

A result like that does not name the mistake. Residual confounding produces it. So would a treatment date recorded one period too early, or an outcome variable built from a window that reaches back past the start. What it does say is that whatever the adjusted models were agreeing about, it was not the effect. They shared an identifying assumption, so they shared its failure. No amount of agreement among them could have surfaced that. Least of all the agreement produced by adding more covariates. Only the estimate run where the answer was known in advance did the job.

The rest of the lesson is about the checks that do that job.

  • A sensitivity parameter is the number you invent to stand in for the bias you cannot measure. You assume a strength for it, then watch what that assumption does to the estimate.
  • A negative-control outcome is one the treatment has no plausible route to. You pick it because the same confounding runs through it anyway. So a treatment that appears to move it has just shown you the confounding. Death before influenza season, at 0.39, is one.
  • A negative-control exposure does the same job from the other end. It stands in for something that could not have caused the outcome, but that was handed out by the same forces that assigned the real treatment.
  • A falsification test does not aim at the effect at all. It aims at something else the causal story implies, so that the story gets a chance to be wrong on its own terms.

Robustness is a claim about which stories you managed to rule out

Sensitivity analysis and falsification ask different questions. The influenza placebo above is what it looks like when only one of them has been asked.

Sensitivity analysis asks how much. Suppose exchangeability is wrong, and the treated and untreated groups differed in something nobody recorded. How strong would that unrecorded difference have to be? How tightly would it have to track the outcome, before the conclusion turns over? That question got a single number in 2017, in Annals of Internal Medicine, and a name. VanderWeele and Ding defined it: “The E-value is defined as the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome to fully explain away a specific treatment-outcome association, conditional on the measured covariates.” Their proposal was that every observational study meant as evidence for causality report an E-value, or some other sensitivity analysis. And that it be computed twice: once for the observed association after adjustment for measured confounders, once for the limit of the confidence interval closest to the null.

It helps to watch the crank turn once. Take Hammond and Horn's smoking data, which Ding and VanderWeele put through the bias factor in 2016. The observed lung-cancer relative risk had a 95% confidence interval of [8.02, 14.36]. Now assume an unmeasured confounder with an exposure-confounder relative risk of 10.73, and a confounder-outcome relative risk of 10.73 as well. That is an enormous hidden variable. The bounding factor is then 5.63, and the corrected interval is [8.02/5.63, 14.36/5.63] = [1.42, 2.55]. Entirely above 1. As they put it: “Thus in fact, not even exposure–confounder and confounder–outcome relative risks of 10.73 suffice to explain away the effect nor the lower confidence limit.” Nothing there was tested. A conclusion was converted into a statement about how big the unseen thing would have to be.

Falsification asks whether. The causal story makes predictions past the main effect, and some of those predictions say a pattern should be absent. So go look for it. If it is there, the story is broken somewhere, and you learn that without ever having to know the true effect. That is precisely how a mortality ratio of 0.39 in a period the vaccine could not reach settled a question eight years of adjusted in-season estimates could not.

Between them they cover the commitments worth writing down before an analysis rather than after. Unmeasured confounding gets a stated magnitude. Timing gets a placebo — an effect looked for before treatment, or outside any window in which treatment could plausibly act. Bias structure gets a negative control, an exposure or an outcome that carries the confounding without the causation. Specification gets stress: other covariates, other windows, other estimators, other targets. The design itself gets a replication attempt on a different data source, a different instrument, a different natural experiment.

What none of this does is work in the other direction. A check that comes back clean has ruled out the threat it was aimed at. About the threats nobody aimed anything at, it has said nothing whatever.

Every one of these checks can convict a design. Not one of them can acquit it.

Example

Each design already tells you what its own falsification test should be

The test is not chosen off a menu. It is read off whatever the design is claiming, which means a design that claims more can be attacked in more places.

It also means the test is itself an estimate, with power of its own. A falsification test that rarely fires is not reassurance. Jonathan Roth made that concrete for difference-in-differences in 2022: “I find that linear violations of parallel trends that would be detected only 50 percent of the time can produce large biases in the treatment effects estimates and lead confidence intervals (CIs) to substantially under-cover the true effect. In the most extreme case, the bias from a trend detected only half the time is larger than the estimated treatment effect and a nominal 95 percent CI contains the true parameter only 24 percent of the time.” He then went looking in the literature. Across three economics journals from 2014 to June 2018 he found 70 papers using an event-study plot, and narrowed those to a replicable sample of 12. Only one of the 12 reported a joint significance test. None discussed what magnitude of pre-trend the data could reject. A flat-looking pre-period plot was doing the work of a test nobody had powered.

  • DiD rests on the two groups drifting in parallel until treatment lands. So look for an effect in the periods before it landed. Then report what your look could actually have caught. In Roth's calibration a trend detected only half the time left a nominal 95 percent CI covering the truth 24 percent of the time — and of the 12 papers he could replicate, only one reported a joint test and none said what magnitude of pre-trend the data could reject.
  • IV rests on the instrument reaching the outcome only by way of the treatment. So point it at outcomes it should have no route to at all. If it predicts those too, it has a second path.
  • RDD rests on units just either side of the cutoff being otherwise alike. So check the things that were fixed before assignment, and check cutoffs where nothing happens. One study of the census of US births from 1983 to 2002 found that one-year mortality falls by about one percentage point as birth weight crosses 1500 grams from above, against a mean infant mortality of 5.5% just above 1500 grams. Barreca and colleagues then ran the same regression discontinuity at 1,000, 1,100, and so on up to 3,000 grams. Most of those are placebo cutoffs, though 1,000 g (extremely low birth weight) and 2,500 g (low birth weight) are real clinical thresholds. The same apparent mortality advantage showed up just below each one, driven by non-random heaping of recorded birth weights at round numbers: “For example, with a bandwidth of 30 grams, 42 of 42 point estimates fall below zero.”
  • Plain observational adjustment rests on having measured the things that matter. So hand the treatment a negative-control outcome it cannot touch, and see whether it moves one anyway. That is the pre-season mortality ratio of 0.39 in the 72,527-person influenza cohort — further from the null than the in-season 0.56 it was supposed to validate.

Analogy

A bridge certified against wind, heat, and vibration

A bridge is certified by loading it. Wind at a stated speed, heat across a stated range, vibration at stated frequencies. It holds, and the certificate goes in the file.

What that certificate records is not that the bridge is safe. It records that the bridge survived those loads at those magnitudes. No earthquake was ever applied, so the file says nothing about one, and everybody involved knows it.

A robustness section works the same way, with one difference that cuts against the analyst. The engineer knows exactly which loads were left off the list. Threats to a causal estimate are partly unobservable, and partly artifacts of whatever model was used to describe them. So the list of untested loads is itself a guess. And unlike the engineer, the analyst can apply a load too gently to matter and still write down that the structure held. That is the pre-trend plot with only a coin's chance of catching the trend it was looking for.

Nobody would sign off on a load certificate with the loads left blank, which is what an unqualified robustness claim is.

Comparison

Agreement, magnitude, and prediction are not the same evidence

Three things get filed under robustness, and they answer three different questions.

Specification stress asks whether the estimate moves when the analyst's arbitrary choices move. Change the covariate list, the window, the estimator, the target, and look. This is the weakest of the three, and the influenza cohort shows why: adjusting for diagnosis-code covariates moved every estimate further from the null rather than closer to it. More adjustment produced more agreement and more bias at the same time, because each adjusted version rested on the same assumption the data could not support.

The hormone-therapy literature is the same lesson at scale. Hormones looked strongly protective of the heart in 59,337 women in the Nurses' Health Study, followed up to 16 years. Grodstein and colleagues reported it in 1996, in the New England Journal of Medicine: “We observed a marked decrease in the risk of major coronary heart disease among women who took estrogen with progestin (multivariate adjusted relative risk, 0.39; 95 percent confidence interval, 0.19 to 0.78) or estrogen alone (relative risk, 0.60; 95 percent confidence interval, 0.43 to 0.83), as compared with women who did not use hormones.” Six years later the Women's Health Initiative randomised the same regimen in 16,608 women. It reported a coronary heart disease hazard ratio of 1.29 (nominal 95% CI 1.02-1.63) over a mean 5.2 years, and was stopped early. A multivariate-adjusted 0.39 and a randomised 1.29 are not two readings of one quantity. They are one design's assumption failing and another design not making it.

Sensitivity analysis asks how much bias it would take. It does not test anything at all. It converts the conclusion into a statement about a threat's size: below this strength the recommendation stands, above it the recommendation does not. The smoking interval surviving a confounder at 10.73 on both arms is one end of that. An E-value small enough that a routine unmeasured difference would be sufficient is the other.

Falsification asks whether a particular prediction of the design holds up. It is the only one of the three that can come back false, which is why it is the only one that puts the analyst at risk. It is also why a pre-season mortality ratio, rather than a pile of agreeing adjusted models, is what found the problem.

Rolling all three into a single robustness score throws away the one distinction that made any of them worth running.

FigureComparison · 3 columns

Sensitivity analysis

How strong must a specified bias be?

  • Quantitative
  • Assumption-dependent
  • Does not detect all biases

Negative control

Does the design generate an effect where none should exist?

  • Can expose confounding
  • Needs credible control
  • May have own pathways

Specification curve

Do conclusions depend on defensible analysis choices?

  • Shows researcher degrees
  • Not all specs valid
  • Still observational

Steps

One row per threat, and no test without a threat attached

Start from the threats, not from the tests you know how to run.

In the first column, write down every way this particular estimate could be wrong: the confounder you suspect nobody measured, the treatment date you are not certain of, the outcome definition that reaches back further than it should, the covariate set another analyst would have chosen instead, the data source you happened to have rather than the one you wanted.

Beside each, write the check that speaks to it and, before running anything, what result would count as bad news. Threats with nothing beside them stay in the table with the cell empty. Those empty cells are the most useful thing in the document. They are the only part a reader cannot reconstruct from the results.

Regulators already require this table in clinical research, and their wording is worth borrowing. The ICH E9(R1) addendum was adopted in November 2019 and came into effect on 30 July 2020. Its glossary defines a sensitivity analysis as “A series of analyses conducted with the intent to explore the robustness of inferences from the main estimator to deviations from its underlying modelling assumptions and limitations in the data.” It requires that sensitivity analyses be pre-specified. It requires that the assumptions underpinning the main estimator be documented. And it requires that each analysis be reported as pre-specified, introduced while blinded, or post hoc. That last requirement is the whole discipline in one line. A check chosen after seeing the result is a different object from the same check chosen before, and the document has to say which one it is.

Then run the checks, record what came back, and write the conclusion in terms of the table instead of in terms of a single word. In the influenza cohort the timing row was filled in, and it came back bad news at 0.39. Had that row been left blank, the in-season 0.56 would have gone out with a covariate list as its only defence. The blank row is where the estimate would have died unnoticed.

FigureProcess · 5 steps
  1. 1

    Name the threat

    Hidden severity, selection, timing, measurement, or interference.

  2. 2

    Choose a test

    Bias parameter, placebo, negative control, or alternative design.

  3. 3

    Predict the pattern

    What should occur if the design is valid or violated?

  4. 4

    Set interpretation

    Failure response and limits of a pass.

  5. 5

    Report jointly

    Primary effect, stress tests, and residual uncertainty.

Key idea

The variable that was already lying around is usually the wrong control

Negative controls tend to get picked from whatever is already in the data, and that is exactly where they go wrong. An outcome chosen because it was to hand may in fact respond to the treatment through a route nobody stopped to consider. A control exposure may feed into the same downstream machinery the real exposure feeds into.

Either way the test stops meaning anything, and it stops meaning anything in both directions at once. A nonzero result no longer points at confounding, because the treatment may genuinely have had an effect there. A null result no longer clears the design, because the control may not carry the confounding in the first place.

The requirement has a published statement and a name. Lipsitch and colleagues set it out in Epidemiology in 2010: “The essential purpose of a negative control is to reproduce a condition that cannot involve the hypothesized causal mechanism, but is very likely to involve the same sources of bias that may have been present in the original association.” The second half of that sentence is what they call U-comparability. The unobserved common causes of exposure A and outcome Y must be as nearly identical as possible to those of A and the negative-control outcome N. Or to those of the negative-control exposure B and Y. Their own worked example is the pre-season influenza analysis, offered as a negative control detecting residual confounding.

So a control earns its place the same way the main analysis did. A graph that shows why the causal arrow is absent. Timing that makes the arrow impossible. Measurement that is not quietly shared with the outcome of interest. And somebody with domain knowledge prepared to say out loud that the treatment cannot reach it.

A control you cannot defend produces a test whose failure and whose success mean the same thing, which is nothing.

The checks are only worth running if they can change what you recommend

All of this matters at one point, which is where somebody acts on the estimate.

The strong case is narrow and can be stated in a sentence. The plausible range of hidden bias was written down before the analysis rather than chosen afterwards to be survivable, the estimate stays on the same side of the decision threshold across that whole range, and the falsification tests the design implied came back the way the design said they would. The smoking illustration is what that looks like when it holds: a confounder at 10.73 on both arms still leaves [1.42, 2.55]. The recommendation stands, and it stands for reasons a reader can inspect.

The weak case looks like the opening one. A modest and entirely believable violation flips the sign, or a placebo fires — a pre-season 0.39 sitting under an in-season 0.56, a hormone regimen at 0.39 observationally and 1.29 in a randomised trial of 16,608 women. When that happens, the honest move is to downgrade the claim inside the document. Say what the decision would need and what evidence could supply it, rather than keeping the number and softening the verb around it.

And retire the bare word. Robust with nothing after it tells a reader that checks were run while hiding which ones. The long version is different. Robust to unmeasured confounding up to the strength we assumed, tested against a pre-treatment placebo whose power to detect a linear violation we report, and never tested against a mis-recorded treatment date. That is far less flattering, and it is the only version anyone can argue with. It is also, for clinical trials since 30 July 2020, close to what the regulator asks for in writing.

A reader cannot check a claim that never says what it survived, and a claim nobody can check is not evidence.

Key takeaways