Causal inference
Identification Versus Estimation
Distinguish causal identification, statistical estimation, finite-sample uncertainty, nuisance modeling, and partial identification.
By the end you can
- Explain identification as a property of assumptions and the data-generating process
- Explain estimation as a finite-sample statistical task
- Recognize why flexible machine learning cannot repair non-identification
- Use bounds or sensitivity analysis when point identification is unavailable
Example
Thirty-two trials emulated with one estimator, and two different answers
When the same machinery is applied thirty-two times and agrees with the truth in one half of the cases and not the other, the machinery is not the thing that varied.
Thirty-two randomised trials, each one reproduced from insurance claims. That was RCT-DUPLICATE, funded by the FDA, and the results ran in JAMA in 2023. Every emulation used the same design template — a propensity-score-matched new-user cohort study — drawn from three US claims databases: Optum Clinformatics, MarketScan and Medicare. The estimator family was held fixed across all 32. What varied was whether the design elements that define the question — population, intervention, comparator, outcome, time — could be emulated in claims data at all.
Overall agreement with the randomised results was Pearson r = 0.82 (95% CI, 0.64-0.91). That is the figure that fits on a slide. Split by design fidelity, it comes apart. The 16 emulations that could reproduce the trial's design elements reached r = 0.93 (0.79-0.97). The 16 where those elements could not be emulated reached r = 0.53 (0.00-0.83). Same databases, same estimator, same investigators. The lower end of that second interval is zero.
The authors put the finding in one sentence: “Real-world evidence studies can reach similar conclusions as RCTs when design and measurements can be closely emulated, but this may be difficult to achieve.” Two different questions were on the table in each of those 32 attempts. Only one of them is a question an estimator can answer.
- The first question is whether the effect you care about can be written down at all, using what the data records, under assumptions someone would defend out loud. In RCT-DUPLICATE that meant the population, intervention, comparator, outcome and time window of the trial being emulated.
- The second is how to compute that quantity well from a finite sample. Algorithms, estimators and standard errors live here. The propensity-score-matched new-user design was common to all 32 emulations, so this half was held constant by construction.
- Where the first question had an answer, agreement with the randomised result was r = 0.93 (0.79-0.97). Where the design elements could not be emulated, it fell to r = 0.53, on an interval running from 0.00 to 0.83.
- The pooled r = 0.82 (0.64-0.91) is a real number, computed correctly. It averages the two halves into a figure that describes neither of them.
Identification comes before algorithms, and regulators write it down
Identification means your assumptions leave a causal quantity exactly one possible value, given the distribution of data you can actually observe. That is a property of the question and the design around it. It is settled before anyone picks a model, and picking a different model does not revisit it.
This is not the house style of one research school. It is a regulatory requirement. The harmonised addendum on estimands and sensitivity analysis, ICH E9(R1), was adopted in 2019, and it makes the separation explicit: “This framework enables proper trial planning that clearly distinguishes between the target of estimation (trial objective, estimand), the method of estimation (estimator), the numerical result (“estimate”, see Glossary), and a sensitivity analysis.” Europe adopted it with legal effect from 30 July 2020. The FDA issued it as final guidance in 2021. Sponsors in the EU, the US and Japan are held to it.
Estimation is the separate job of approximating that identified quantity from a finite sample. This is where flexible methods earn their reputation. They fit the auxiliary pieces better. They pick up differences between subgroups. They lose less to a badly guessed functional form. All of that is real, and all of it happens downstream of the target being fixed.
What no learner can do is manufacture information that was never recorded. If treatment was assigned partly on a judgement nobody wrote down, that judgement is absent from every row, in every fold, at every sample size. If no untreated person resembles the treated ones, fitting harder does not invent one. A finer ruler gives you a more precise reading of whatever point you decided to measure. It has nothing to say about whether the point belongs where you put it. The comparison is unfair in one direction: a ruler is fixed, while causal assumptions can be argued with, probed against a design, sometimes falsified outright. But that argument is made in sentences, not inside the fit.
A learner can only rearrange information that is already in the table, which is why ICH E9(R1) fixes the target of estimation before the table exists.
Comparison
Two failures that look the same on the screen
An unidentified effect and a badly estimated one arrive looking alike: a number, an interval, a plot with a few odd points. The remedies have nothing in common.
The cleanest documented pair uses one dataset twice. The Nurses' Health Study data on postmenopausal hormone therapy was re-analysed in 2008 by Hernán and colleagues, in Epidemiology. No new data was collected. No exotic estimator was reached for. They re-specified the question as a sequence of emulated trials, with a defined eligibility, a defined time zero and an intention-to-treat contrast, and reported that “The ITT hazard ratios (HRs) (95% confidence intervals) of CHD for initiators versus noninitiators were 1.42 (0.92-2.20) for the first 2 years, and 0.96 (0.78-1.18) for the entire follow-up.” That 1.42 points the same way as the hazard ratio of 1.29 the randomised Women's Health Initiative had found in 2002 — out of a cohort whose earlier analyses had shown a strong protective association. What moved the answer was eligibility, time zero and the contrast. Three sentences of design, not one line of code.
Estimation error responds to work you can do at your desk: a better nuisance model, a more robust estimator, more data of the kind you already have. Design failure responds to none of it. It asks for a different question, a different population, a different way of generating the data, or a frank statement of what cannot be recovered. Reaching for the first set of tools when the fault lies in the second is how a team ends up steadily improving a number that was never measuring the treatment.
Non-identification
Several causal worlds fit the same observed data.
- Needs assumptions or new design
- Cannot be fixed by sample size
- May require bounds
Model misspecification
Chosen nuisance or outcome model is wrong.
- Use diagnostics and flexible models
- Cross-fit where appropriate
- Compare estimators
Sampling variability
Finite samples produce uncertainty.
- Use valid inference
- Increase effective information
- Report interval or distribution
Example
Identification failures disguised as modelling tasks
The 16 emulations that landed at r = 0.53 have a family. In each of the following the work looks like a modelling task, the software runs without complaint, and what is missing is causal information that no estimator supplies.
- No overlap: when the highest-risk patients are all treated, the question of what would have happened to them untreated has nobody in the data to point at. The model answers from the milder cases it did see. D'Amour and colleagues showed in 2021 that this gets worse as covariates accumulate. Their Theorem 1 bounds the Euclidean distance between the treated and control covariate mean vectors by min{||Σ0||_op^(1/2)·B^(1/2), ||Σ1||_op^(1/2)·B^(1/2)}, with constants that depend only on the overlap bound and are free of p; Corollary 1 multiplies that bound by p^(-1/2) for the average per-covariate discrepancy. In high dimensions, strict overlap forces the covariate means towards exact balance.
- Hidden confounding: 59,337 women, followed by the Nurses' Health Study for up to 16 years. Its investigators reported in 1996, in the New England Journal of Medicine, that “We observed a marked decrease in the risk of major coronary heart disease among women who took estrogen with progestin (multivariate adjusted relative risk, 0.39; 95 percent confidence interval, 0.19 to 0.78) or estrogen alone (relative risk, 0.60; 95 percent confidence interval, 0.43 to 0.83), as compared with women who did not use hormones”. The randomised Women's Health Initiative, 16,608 women, stopped early on 31 May 2002 after a mean 5.2 years. It found a CHD hazard ratio of 1.29 (nominal 95% CI, 1.02-1.63) on 286 cases. Precision was never the problem: 0.39 and 1.29 are on opposite sides of 1.
- Post-treatment selection: in 1991 US births, infants of smokers had higher low-birth-weight risk and higher infant mortality overall. Yet among low-birth-weight infants, mortality was lower for infants of smokers — relative rate 0.79. Hernández-Díaz and colleagues took that paradox apart in 2006. As their abstract puts it, “The authors use causal diagrams to show that, even in the absence of any beneficial effect of smoking, an inverse association due to stratification on birth weight can be found.” A better estimator would have estimated 0.79 more precisely.
- Ill-defined intervention: one exposure flag switched on for several versions of a treatment that behave differently from each other. The estimate is then an average over a mixture nobody would choose to deploy. That is why intervention and comparator sit in the list of design elements RCT-DUPLICATE had to reproduce before its emulations agreed with the trials.
Visual
Four layers sit under every causal number
Any causal result handed to you can be pulled apart into four layers. ICH E9(R1) gives three of them names.
At the top is the estimand, defined in the addendum as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective”: for whom, under which intervention, compared with what. Below it is the argument that turns that quantity into something computable from observed data, together with the assumptions it leans on. Below that is the estimator, “A method of analysis to compute an estimate of the estimand using clinical trial data”. At the bottom is the estimate, “A numerical value computed by an estimator”, and its reported uncertainty.
Read a result from the bottom and the printed number is all you ever see. Read it from the top and you find out whether the quantity was reachable at all. The 16 RCT-DUPLICATE emulations that reached r = 0.53 did not fail at the estimator layer. They failed one layer above it, where the population, intervention, comparator, outcome and time could not be reconstructed.
The addendum is explicit about the order in which the layers are built: “The targets of estimation are to be defined in advance of a clinical trial.”
- 01
Causal estimand
Counterfactual quantity required by the decision.
- 02
Identification formula
Observed-data functional under assumptions.
- 03
Estimator
Procedure applied to a finite sample.
- 04
Estimate and uncertainty
Numerical result with sampling and modeling error.
Steps
Write the identification argument before the estimator
Take an analysis someone is proposing and force the two halves apart on paper, in that order.
The first half is prose and mentions no software. It names the population, the intervention and the comparison. It states the assumptions that would let observed data stand in for outcomes nobody got to see. It says where those assumptions would break. If this paragraph cannot be written, there is nothing for an estimator to estimate.
That has been turned into a published checklist. In 2025 Imbens and Xu revisited LaLonde's 1986 experimental-benchmark critique, running modern estimators on the same data. Their finding is a warning about the second half: “We show that modern methods, when applied in contexts with sufficient covariate overlap, yield robust estimates for the adjusted differences between the treatment and control groups. However, this does not imply that these estimates are causally interpretable.” Robustness across methods is not evidence of identification. What they recommend instead is examining the assignment process, inspecting overlap and running placebo tests. Goodness-of-fit tests they judge inadequate for the job.
The second half is the numerical procedure. It becomes worth arguing about only once the first half survives being read by someone hostile. Written in this order, a missing comparison group turns up in a sentence, before anyone fits anything.
- 1
State the estimand
Write the target counterfactual contrast.
- 2
List assumptions
Consistency, exchangeability, positivity, interference, measurement.
- 3
Derive the functional
Show which observed quantities identify the target.
- 4
Choose the estimator
Regression, weighting, matching, g-method, or design estimator.
- 5
Plan alternatives
Bounds, sensitivity analysis, redesign, or partial identification.
Example
Four words that keep the two halves apart
The whole distinction rests on these, so it is worth being exact about what each one means.
- Identification is the claim that your assumptions leave exactly one possible value for the causal quantity given the observable distribution. It is a statement about logic, made before any sample exists. It is the layer ICH E9(R1) calls the target of estimation.
- Estimation is everything that happens after that, when a finite sample is used to approximate the value and uncertainty enters the picture. In the addendum's vocabulary, the estimator is “A method of analysis to compute an estimate of the estimand using clinical trial data”.
- A nuisance function is one of the auxiliary pieces an estimator needs along the way, such as a regression of the outcome or a model of who got treated. Getting them right improves the estimate. It does nothing for the argument above it.
- Partial identification is the honest fallback when the assumptions support a range rather than a single value. Manski and Molinari's COVID-19 bounds of [0.004, 0.525] for Illinois are a published example. Reporting that range counts as a result, not a failure to produce one.
Key idea
Machine learning makes an empty region look smooth
A flexible learner will return predictions where there was little or no treatment support. It does not refuse and it does not flag. Its surface passes through the empty region as smoothly as through the crowded one, because smoothness is what it was built to produce.
The empty region is not a sign of carelessness. In high dimensions it is close to arithmetically unavoidable. The trade-off is stated in the first two sentences of D'Amour and colleagues' abstract: “Researchers often argue that unconfoundedness is more plausible when more covariates are included in the analysis. Less discussed is the fact that covariate overlap is more difficult to satisfy in this setting.” Their Theorem 1 bounds the distance between the treated and control covariate mean vectors by a constant free of p, and Corollary 1 scales the per-covariate version by p^(-1/2). Whenever the covariance operator norms grow more slowly than p, as they do for independent or weakly dependent covariates, the bound tends to zero. Every covariate you add to make unconfoundedness defensible makes strict overlap harder to satisfy.
Cross-validation does not catch this. Holding rows out tests predictions of outcomes that were observed. The outcome that is missing in the unsupported region is missing in every fold as well. Excellent validation numbers are entirely consistent with a fabricated counterfactual.
What does catch it is unglamorous, and it is what Imbens and Xu recommend. Look at where treatment support actually lies. Ask which few rows the estimate leans on. Write the assumptions out in words. Run a placebo test. Then try to tell a rival causal story that fits the same data equally well. The sophistication of the method is not one of these checks and cannot substitute for any of them.
Flexibility reduces the errors that come from having a finite sample; strict overlap failing at rate p^(-1/2) is not that kind of error.
When to stop estimating and say so
If the identification argument rests on assumptions you would not defend to a hostile reader, or on a population where the comparison does not exist, more estimation is not a remedy. It is camouflage. Complexity buys the appearance of rigour and moves the weak joint further from view.
Four honest moves remain, and none of them is a better estimator. Narrow the target to the part of the population where a comparison genuinely exists. Collect the data that would record whatever is driving treatment. Find or build a design that carries the argument on its own — the route that took the Nurses' Health Study data from a strong protective association to 1.42 (0.92-2.20) in the first 2 years. Or report the set of values the assumptions actually support, and say what it would take to shrink it.
The fourth move is not a theoretical option. In the spring of 2020 nobody could say how many people had been infected, because test data was missing non-randomly. Manski and Molinari published the range instead of a number: “We find that the cumulative infection rates as of April 24, 2020, for Illinois, New York, and Italy are, respectively, bounded in the intervals [0.004, 0.525], [0.017, 0.618], and [0.006, 0.471].” The corresponding infection-fatality rates were bounded to [0, 0.033], [0.001, 0.049] and [0.001, 0.077], against reported death rates among confirmed cases of 0.045, 0.059 and 0.134. Those intervals are wide enough to be uncomfortable. They are also what the data and the stated assumptions support. A bounded claim over the region with real comparisons, plus a stated blank for the rest, is a smaller result than a point estimate. It is a true one.
Once identification is credible, the choice of estimator becomes worth arguing about on its own terms: robustness, efficiency, the diagnostics it exposes, and whether it fits the design that made the argument work in the first place.
A report that folds its causal argument inside its estimator has answered a question no reader can check.
Key takeaways
- Identification is logically prior to estimation, and no choice of algorithm reopens it — ICH E9(R1) requires the estimand to be defined in advance of the trial.
- Whether an effect is identified depends on the assumptions you will defend and on the design elements the data can reproduce: population, intervention, comparator, outcome, time.
- Holding the estimator fixed and varying only design fidelity moved RCT-DUPLICATE's agreement with 32 randomised trials from r = 0.93 (0.79-0.97) to r = 0.53 (0.00-0.83).
- Predictive flexibility cannot repair hidden confounding — 59,337 women and 16 years of follow-up gave 0.39 where the randomised trial gave 1.29 — nor a region where strict overlap fails.
- A range the assumptions genuinely support beats a single number that holds only if they do; Manski and Molinari published [0.004, 0.525] rather than a point estimate.
- Show the identification argument separately from the estimator, because robust agreement across modern methods “does not imply that these estimates are causally interpretable”.