Causal inference
The G-Formula and Standardization
Derive the adjustment formula, distinguish conditional and marginal effects, and diagnose extrapolation in standardization.
By the end you can
- Derive the point-treatment g-formula under exchangeability and positivity
- Distinguish conditional outcome models from marginal causal estimands
- Standardize predicted outcomes to a declared target population
- Diagnose unsupported counterfactual prediction and model dependence
Example
They reported a coefficient and called it the policy effect
The whole mistake runs in public, on a dataset anyone can download. NHEFS smoking cessation: 1,566 people, quitting smoking as the treatment, weight change in kilograms as the outcome. Fit one outcome model, with a single product term between quitting and smoking intensity. The coefficient sitting next to the treatment prints as 2.6 kg. Hernán and Robins work exactly this example in Causal Inference: What If.
Read that number off, call it the effect of quitting on weight, and you have reported something correct that nobody asked for. The product term is in the model. So quitting does different things to people who smoke different amounts. The conditional effect is 2.8 kg (95% CI 1.5, 4.1) at 5 cigarettes a day and 4.4 kg (95% CI 2.8, 6.1) at 40. The 2.6 kg is a comparison made at fixed covariate values, between people already alike. Nobody in the room had asked about that comparison. The question was about the 1,566 as they actually are.
The number that answers it is 3.5 kg (95% CI 2.5, 4.5). It comes out of the same fitted model, on the same people, in four moves rather than one.
- The outcome model gives you one thing: the expected weight change for a person with a given treatment and a given set of covariates. Be strict about how little that is. With the product term in, that conditional quantity is 2.8 kg at 5 cigarettes a day and 4.4 kg at 40. It is not one number at all.
- From it you predict every unit in the target twice over: all 1,566 as though each of them had quit, then all 1,566 again as though each had carried on smoking.
- You average each set of predictions across the target covariate distribution — 5.2 kg under quitting, 1.7 kg under continuing. That averaging is standardization. Nothing more mysterious than that.
- The contrast between the two standardized means is 3.5 kg (95% CI 2.5, 4.5). That is the marginal effect, and it is the number the policy question was asking for all along — not the 2.6 kg printed beside the treatment.
What the g-formula lets you claim, and what it makes you assume
Three assumptions do the work, and they are worth naming before the formula rather than after it. Consistency says the outcome you observed under the treatment someone actually received is the outcome they would have had under that treatment. Conditional exchangeability says that once the covariates are held fixed, who ended up treated tells you nothing more about how they would have fared. Positivity says every covariate pattern in your target population had a real chance of receiving each treatment.
Grant those three and the g-formula follows. The mean potential outcome under treatment a equals the average of your conditional outcome model, evaluated at that treatment, over the covariate distribution of the population you care about.
Notice what it is. It is an identification result: a statement about which observable quantity equals the causal one you cannot see. It is not a promise that your estimate is any good. Whatever the outcome model gets wrong — nonlinearities, interactions, treatment-covariate combinations it barely saw — walks straight into the average and comes out the other side wearing the word effect.
Reprice every customer's basket under one fee rule, then under the other, and average the totals across the same customer mix. The gap between the two totals is what switching rules does to that mix, not to any one shopper. Baskets are the generous version of the problem, because the second total can be computed exactly. On NHEFS it cannot. Each of the 1,566 people was observed under exactly one of quitting and continuing. So every one of them contributes a prediction that was never checked against anything, to one of the two standardized means.
The formula tells you which quantity to compute; the three assumptions decide whether the result has earned the word effect.
Example
Four words that are easy to blur and expensive to swap
Four words mark four different places in that move, and they were not all coined at once. Swapping any two of them is how an analysis ends up claiming more than it estimated.
- The g-formula is the identification result itself. Robins published it in 1986, in a 120-page paper on the healthy worker survivor effect that has since collected 2,852 citations. He did not call it the g-formula. A margin note in Causal Inference: What If explains what happened next: “Robins (1986) described the generalization of standardization to time-varying treatments and confounders, and named it the g-computation algorithm formula. Because this name is very long, some authors have abbreviated it to g-formula and other authors to g-computation.”
- Standardization is the arithmetic that result asks for. It is the older move the 1986 paper generalised: averaging conditional predictions over whichever covariate distribution defines the population you named as the target. It is what turns a fitted model into 5.2 kg and 1.7 kg.
- A marginal effect is what comes out of the far end — a population-level contrast between interventions that no longer holds any covariate fixed. It is the 3.5 kg. It is quoted at no particular number of cigarettes a day, because it is not conditional on one.
- Counterfactual extrapolation is the failure hiding inside the prediction step: a confident number produced for a unit under a treatment that units like it almost never received.
Example
The same move under four names
Once you recognise the move — predict twice, average, contrast — you start finding it under names that sound like separate methods and are not. What differs between them is the outcome model, the target population, or both. Each of the four below has a dated, published instance with its numbers on the record.
- Parametric g-computation is the plain version. The outcome model is a regression whose form you specified yourself and can be held responsible for. Its first large-scale epidemiologic application ran on Nurses' Health Study data and was published in December 2009 by Taubman, Robins and two co-authors. The abstract reports: “Over the period 1982-2002, the 20-year risk of CHD in this cohort was 3.50%. Under a joint intervention of no smoking, increased exercise, improved diet, moderate alcohol consumption and reduced body mass index, the estimated risk was 1.89% (95% confidence interval: 1.46-2.41).” The 1.89% is a population-level risk under a policy nobody ran. It is not a coefficient.
- The machine-learning g-formula swaps that regression for flexible nuisance learners, with cross-fitting where the estimator needs it to stay honest. Cross-fitting is not folklore. It is a named construction from the 2018 double/debiased machine learning paper, by Chernozhukov and six co-authors, and the abstract names it in one line: “In order to avoid overfitting, our construction also makes use of the K-fold sample splitting, which we call cross-fitting.” The seventh of those authors is James Robins, who wrote the 1986 g-formula paper.
- Risk standardization holds one case mix fixed for every intervention being compared, so that differences in who was treated stop doing the work the treatment is being credited with. In the United States it is law. 42 CFR § 412.152 tells the Centers for Medicare & Medicaid Services that “Excess readmissions ratio is a hospital-specific ratio for each applicable condition for an applicable period, which is the ratio (but not less than 1.0) of risk-adjusted readmissions based on actual readmissions for an applicable hospital for each applicable condition to the risk-adjusted expected readmissions for the applicable hospital for the applicable condition.” The floor adjustment factor was 0.99 for FY 2013, 0.98 for FY 2014 and 0.97 for FY 2015 onward — a maximum 3% cut to base operating DRG payments. The AMI measure behind it was published by Krumholz and colleagues in 2011. It was built on a 2006 Medicare cohort with an unadjusted 30-day readmission rate of 18.9%, a final model of 31 variables and a C statistic of 0.63.
- Transported standardization takes conditional effects estimated in one population and averages them over the covariate mix of a different one. Same arithmetic, pointed somewhere new. ACTG 320 enrolled 1,156 HIV-infected US adults in 1996: 577 randomised to HAART, 579 to a largely ineffective combination, followed 52 weeks. In 2010 Cole and Stuart standardized that trial to the CDC's estimate of the US population infected with HIV in 2006. Their Results section: “When we simultaneously accounted for differences in age, sex, and race/ethnicity between the trial sample and target population, the hazard ratio was weakened from 0.51 to 0.57.” The intent-to-treat limits were 95% CL 0.33, 0.77. Reweighted they were 0.33, 1.00, the confidence-limit ratio widening from 2.33 to 3.03. Weighting on age alone moved the estimate to 0.68, on sex alone 0.53, on race alone 0.46.
Key idea
The model will answer a question nobody ever ran
Your outcome model will answer anything you put to it. Ask it for the untreated outcome of a unit that is treated in nearly every case it has ever seen, and it returns a number to as many decimal places as you like, with no warning attached. Ask it for the treated outcome of someone who was never eligible for treatment, and it does the same.
The failure has a clinical description. Petersen and four colleagues published it in 2012, and the abstract states it without hedging: “Positivity violations occur when certain subgroups in a sample rarely or never receive some treatments of interest. The resulting sparsity in the data may increase bias with or without an increase in variance and can threaten valid inference.” A smooth prediction is not an empirical comparison. Where the data thinned out, the model filled the gap from its functional form. That filling is what you are about to average.
The same paper supplies the discipline, and it is worth following rather than improvising. Their diagnostic is the parametric bootstrap, run to gauge how severe a positivity violation actually is. Their responses are four named moves: restriction of the covariate adjustment set, use of an alternative projection function to define the target parameter within a marginal structural working model, restriction of the sample, and modification of the target intervention.
So look before you average. Map where treatment actually varies across the covariate space. Display the parts of the target that rest on extrapolation, instead of letting them disappear into the total. Restrict the target population where a region has nothing to compare against. Then estimate the same contrast by weighting or by a design-based route and put the two side by side. When they disagree, the disagreement is usually sitting in exactly the thin region you were hoping to ignore. And keep in view what averaging does to variety. The AMI measure ran 3,890 hospitals into risk-standardized readmission rates. The 25th and 75th percentiles were 18.6% and 19.1%, the 5th and 95th 18.0% and 19.9%. Thousands of hospitals, one narrow band.
Averaging is what makes extrapolation invisible — the thin regions dissolve into one tidy number that looks like all the others.
Steps
Pull the four moves apart by hand, once
In most software those four moves collapse into a single command. Which is precisely why they are worth separating by hand once, on a dataset small enough to print. NHEFS is the obvious candidate, because someone has already published the numbers you should get. Tom Palmer's R and Stata companion to Causal Inference: What If prints them.
Fit the outcome model — weight change on quitting, on smoking intensity, and on the product term between them. The coefficient on quitting comes out as 2.5595941 and the product term as 0.0466628. Then stop. Do not read the first of those out loud yet.
Predict the entire dataset as though every one of the 1,566 had quit. Predict it again as though none had. Average each column of predictions on its own: 5.178841 under quitting, 1.660267 under continuing. Only then take the difference, 3.518574 — the 3.5 kg. For contrast, the conditional route on the same fit gives 2.7929 at 5 cigarettes a day and 4.4261 at 40, and never gives 3.5 anywhere.
Doing it in that order makes something obvious that no amount of arguing will. The treated average is not an average over treated units. It is an average over everybody, priced under treatment. That is the claim the 2.5595941 was never in a position to make.
- 1
Fit E[Y|A,X]
Include clinically or substantively justified terms.
- 2
Duplicate target rows
Create one copy per intervention strategy.
- 3
Generate predictions
Use the same fitted model for both strategies.
- 4
Average by strategy
Preserve the declared target covariate distribution.
- 5
Quantify uncertainty
Bootstrap units or use valid influence-based inference.
Visual
The order of the workflow is doing work
The workflow is short enough to hold in your head, and its order is not arbitrary. You define the target first, because every quantity downstream is an average over it. You fit the outcome model. You predict under treatment, then under the comparison, using the same units both times. You average and contrast.
Follow it on the map with the assumptions written next to the stages they actually govern. Exchangeability is a claim about the fitting step. Positivity is a claim about the two prediction steps, and it is the one that quietly fails. That is why it, and not the others, has a diagnostic literature of its own. The choice of target is not an estimate at all. It is a decision, and it belongs to whoever asked the question rather than to whoever fitted the model.
- 1
Define target
Choose population, treatment strategies, outcome, and scale.
- 2
Fit outcome model
Estimate conditional mean using pre-treatment covariates.
- 3
Predict under treatment
Set A=1 for every target unit.
- 4
Predict under comparison
Set A=0 for every target unit.
- 5
Average and contrast
Compute marginal intervention means and uncertainty.
Comparison
Both can be right and still disagree
A conditional quantity answers what happens among units that match on the covariates. A marginal quantity answers what happens to a population with a particular mix of them. Different questions, so different answers, and no contradiction anywhere in sight. On NHEFS both come out of one fit: 2.6 kg beside the treatment, 3.5 kg (95% CI 2.5, 4.5) after standardization, and 2.8 kg and 4.4 kg at 5 and 40 cigarettes a day.
Causal Inference: What If marks the fork explicitly when it reaches it. Its chapter on outcome regression notes: “In Section 13.2, outcome regression was an intermediate step towards the estimation of a standardized outcome mean. Here, outcome regression is the end of the procedure. Rather than standardizing the estimates of the conditional means to estimate a marginal mean, we just compare the conditional mean estimates.” Same model, two destinations. The software will take you to either without comment.
The two coincide only when several things hold at once: a model linear in the right way, no interaction between treatment and covariates, an effect that is the same for everyone. The product term between quitting and smoking intensity breaks the second of those on its own. The coefficient then drifts away from the standardized contrast while both remain correctly estimated.
Which means choosing between them is not a statistical decision. It follows from who will act on the number. Someone advising one person who smokes 40 a day wants 4.4 kg. Someone deciding whether to run a cessation policy across the whole population wants 3.5 kg. Handing them 2.6 kg instead is the error the opening analysis made.
Conditional association
Compares treatment at fixed covariates.
- Model-specific
- May be noncollapsible
- Not automatically causal
Conditional causal effect
Contrasts potential outcomes within covariate strata.
- Can vary by X
- Requires identification
- Useful for heterogeneity
Marginal causal effect
Averages over a target covariate distribution.
- Population-level
- Decision-friendly
- Changes with target mix
When to believe a standardized effect
Standardization earns trust in a narrow band of conditions. The outcome model has to be estimable across the regions you are averaging over, not merely fittable on the sample as a whole. The target population has to be stated out loud, rather than inherited by accident from whoever assembled the data.
That second requirement is no longer advice. A standards body has written it down. ICH E9(R1), the Addendum on Estimands and Sensitivity Analysis in Clinical Trials, was adopted on 20 November 2019, after a public consultation that opened on 30 August 2017. It requires the treatment effect to be built from named attributes before any estimator is chosen: the treatment condition, the population of patients targeted by the clinical question, the variable, the handling of intercurrent events, and a population-level summary. Its glossary defines an estimand as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective. It summarises at a population-level what the outcomes would be in the same patients under different treatment conditions being compared.” The same patients under different treatment conditions is the predict-twice step, written out by a standards body. The FDA issued the guideline as final guidance for industry on 12 May 2021.
Flexible learners help with the first requirement, since they reduce the damage a wrong functional form can do. They buy you nothing at all on support. So they arrive with obligations attached: check where treatment is supported, check that predictions are calibrated, and refit under different specifications to see how far the answer travels.
When the answer turns out to rest on unsupported tails, there are two honest exits, and neither of them is trusting the model to be smooth. Shrink the estimand to the population you can actually speak about — restriction of the sample, in the positivity paper's taxonomy. Or change what you are asking of it, which is their modification of the target intervention. They are candid about the cost: their responses “can be understood as trading off proximity to the initial target of inference for identifiability”. The third exit is to go back and get a stronger design.
A standardized effect is worth no more than its thinnest region, however good the overall fit looks.
Key takeaways
- The g-formula identifies the mean outcome under an intervention, provided consistency, conditional exchangeability and positivity hold. Robins published it in 1986, under the name g-computation algorithm formula.
- A conditional regression coefficient and a marginal causal effect are different quantities, and a correct coefficient can still be the wrong answer: on one NHEFS fit of 1,566 people, 2.6 kg sits beside the treatment while the standardized contrast is 3.5 kg (95% CI 2.5, 4.5).
- Standardization means predicting every unit under each treatment and averaging those predictions over the population you named as the target — 5.2 kg if all 1,566 quit, 1.7 kg if none of them do.
- When the treatment does different things to different people, the outcome model has to carry treatment-covariate interactions: with the product term in, the conditional effect is 2.8 kg (95% CI 1.5, 4.1) at 5 cigarettes a day and 4.4 kg (95% CI 2.8, 6.1) at 40.
- A model that fits observed outcomes well has said nothing about the counterfactual predictions nobody was ever able to check. That is why Petersen and colleagues pair a parametric-bootstrap diagnostic with four named retreats, rather than a smoother fit.
- Support diagnostics and the named target belong next to the standardized estimate, not in an appendix behind it. ICH E9(R1), adopted 20 November 2019, makes the population and the population-level summary part of the estimand itself.