Causal inference
Marginal Structural Models, G-Computation, and G-Estimation
Compare inverse probability treatment weighting, parametric g-formula, and structural nested models for longitudinal effects.
By the end you can
- Explain how marginal structural models use treatment and censoring weights
- Describe longitudinal g-computation for static and dynamic strategies
- Explain the role of structural nested models and g-estimation
- Choose diagnostics for treatment history support and weight stability
Example
The estimate fit cleanly, and a handful of patients were holding it up
918 HIV-infected US men and women, followed a median of 5.8 years between 1996 and 2005. Nobody was randomised. The people who kept taking the drug were not like the people who stopped. So each person was reweighted by how surprising their own sequence of decisions looked — surprising given what had been recorded about them at the time each decision was made. Cole and Hernán published that analysis in 2008, with the numbers attached.
The untruncated stabilised weights had a mean of 1.04 and a standard deviation of 1.15. Their maximum was 46.8.
That is the whole lesson in three numbers. The typical person in that analysis counted about once. Somebody in it counted 46.8 times. The weighting model considered the history they had actually lived very unlikely, and the inverse of a small probability is a large number. The model fits cleanly either way. Nothing in its output complains. What the output does not say is that the final estimate is substantially the story of the few people at the top of that distribution. The analysis has built a pseudo-population: a reweighted copy of the data in which treatment no longer depends on the past. The copy can be perfectly well behaved and nearly empty of information.
Four terms do the work for the rest of this lesson. The case above is easier to read once they are separate.
- A marginal structural model is a model of the outcome you would have seen under one whole treatment history rather than another. It is fitted on the reweighted copy, not on the raw data.
- The weight carried by a patient is the product of the inverse probabilities of each treatment decision they actually made — one factor for every point where the decision could have gone the other way. That is how a mean of 1.04 and a maximum of 46.8 end up in the same column.
- G-computation refuses the reweighting entirely. It models how covariates and outcomes evolve over time, then simulates the population forward under whichever strategy you want to ask about.
- G-estimation starts from the residuals of treatment assignment — the part of what people did that the recorded history failed to predict — and uses them to recover the parameters of a structural nested model.
Example
Checks that pass at one decision can say nothing about a sequence
The usual checks are computed one decision at a time. Is the treated group comparable to the untreated group at this moment? Does everyone here have some chance of either assignment? Every one of those checks can come back clean while the product of them falls apart. Probabilities multiply. Comparisons do not.
Treating that as a real problem rather than a box to tick has a protocol behind it. Petersen and colleagues wrote it in 2012, and their statement of the problem is flat and worth having in front of you: “Positivity violations occur when certain subgroups in a sample rarely or never receive some treatments of interest. The resulting sparsity in the data may increase bias with or without an increase in variance and can threaten valid inference.”
Note what that sentence does not say. It does not say the data violate positivity, full stop. Positivity is set out there as a model- and parameter-specific identifiability problem — a property of the estimator and the question, not a defect the dataset carries around on its own. The paper proposes the parametric bootstrap as a diagnostic for how badly sparsity is biasing a given estimator. It then catalogues four responses. Each of the four buys identifiability by moving away from the question you started with. The honest version of the work is to say which move you made.
- Restrict the adjustment set. Carry fewer covariates, so that fewer thin combinations of prior treatment and history exist to be conditioned on — at the cost of the confounding control those covariates were there to provide.
- Change the projection function of the marginal structural working model. Keep the data and change the shape of the summary you are extrapolating over, so the model borrows strength across strata rather than leaning on the sparse ones.
- Restrict the sample. Analyse only the people whose histories the strategies could plausibly have produced, and accept that the answer now describes that subgroup and not the population you began with.
- Modify the target intervention. Ask about a strategy the data actually support instead of the one you first wrote down — the most explicit of the four, because it changes the question in the open.
Three ways out of the same trap, each paying for it somewhere else
The trap is worth stating plainly, because all three methods exist because of it. In a study that follows people over time, one variable can be both a consequence of earlier treatment and a cause of the next treatment decision. Adjust for it the ordinary way and you erase the part of the effect that travelled through it. Leave it out and it confounds every later decision. Ordinary regression has no setting for a variable that is both.
Zidovudine and survival in the Multicenter AIDS Cohort Study showed what that costs. The founding applied paper on marginal structural models ran that analysis in Epidemiology in 2000. CD4 count sits in exactly the double role: a confounder of the next treatment decision, and a consequence of past treatment. The crude mortality rate ratio was 3.6 (3.0–4.3). Adjustment for baseline covariates by standard methods brought it to 2.3 (1.9–2.8) — still, read at face value, a drug that kills the people who take it. Then the abstract: “Using a marginal structural Cox model to control further for time-dependent confounding due to CD4 count and other time-dependent covariates, the mortality rate ratio was 0.7 (95% conservative confidence interval = 0.6-1.0).” One dataset, one exposure, three numbers: 3.6, then 2.3, then 0.7. The estimate crosses from strong harm to apparent benefit purely by changing how the sequence of decisions is handled.
Marginal structural models take the first route out. They leave the confounding out of the model and push it into the weights, reweighting the observed histories until treatment no longer depends on the past — the pseudo-population from the case.
The parametric g-formula takes the second. Rather than reweighting anyone, it models the whole longitudinal process: how covariates move, how outcomes follow. Then it simulates that process forward under an intervention you name.
Structural nested models take the third. They parameterise the blip: the small change in outcome from a bit more treatment at one moment, given everything that came before. G-estimation then finds those parameters through treatment residuals.
Picture travellers moving through a sequence of junctions, where the sign at each junction was reset by whoever came through earlier. You can reweight the travellers you actually observed until the signs stop mattering. Or you can model the signs and send simulated travellers through instead. Either way, a rare complete journey may rest on very few people. Junctions at least stand still. A patient's choices can also turn on how they felt that week. Nobody wrote that down, and no amount of reweighting or simulation recovers what was never recorded.
What all three share is the price of admission. At every decision point, the recorded history has to be enough to make treatment as good as random. And the histories you want to compare have to be histories somebody actually lived.
None of the three buys an answer without assumptions; they differ only in which model you end up betting the result on.
Key idea
Every decision was ordinary and the whole sequence was still rare
Take a patient whose treatment at each visit was, say, moderately predictable — not a surprise, just not a certainty either. Multiply that across a long follow-up and the probability of the exact sequence they lived becomes tiny. The weight is the inverse of that product, so it becomes enormous. Nothing unusual happened at any single visit. The rarity is entirely a property of the sequence. In those 918 people it is the distance between a mean stabilised weight of 1.04, a standard deviation of 1.15, and a maximum of 46.8.
The standard repairs are stabilisation and truncation, and both work. Cole and Hernán are candid about the status of the second one: “Weight truncation is presented as an informal and easily implemented method to deal with these tradeoffs.” Informal, and consequential. Their published ladder walks the trade in six rows. With untruncated stabilised weights the estimated coefficient was −1.91, with a standard error of 0.132. Truncating at the 1st and 99th percentiles moved it to −1.80 (SE 0.122). At the 5th and 95th, −1.73 (SE 0.106). At the 10th and 90th, −1.69 (SE 0.101). At the 25th and 75th, −1.63 (SE 0.091). Truncating at the median, −1.59 (SE 0.089).
Read the two columns together. The standard error falls monotonically, by about a third from top to bottom, so the estimate looks steadily more precise. Over the same six rows the point estimate drifts 0.32 units in one direction. Cutting the extreme weights off cuts away part of the correction they were making. What is left behind is residual confounding, underneath a number that now looks reassuringly stable. Nothing in the output flags which of the six rows you should have reported.
That trade-off is a judgement, so it should be visible rather than buried in a footnote. Show how the weights evolve over time. Report the effective sample size the weights actually deliver. State what fraction of observed histories is compatible with each strategy. And do what Cole and Hernán simply did: print the whole ladder rather than the rung you liked.
Positivity here is a claim about entire journeys, and no check performed one step at a time can vouch for it.
Visual
The order of the work, before any model is fitted
Most of this is decided before there is code. Say first what strategy you are actually asking about: stay on therapy throughout, stop at a threshold, follow a rule that responds to the patient's state. Then decide how history will be represented. That means choosing what counts as the past at each decision point, and what gets carried forward. Only then choose the g-method. That choice depends on which parts of the system you can model credibly, and you do not know that until the first two questions are answered.
This ordering is not one author's preferred habit. It is written into an adopted international guideline. ICH E9(R1), the harmonised addendum on estimands and sensitivity analysis, was adopted at Step 4 in November 2019 and has had effect in the EU since 30 July 2020. On constructing an estimand it puts it this way: “it is important to proceed sequentially from the trial objective and an understanding of the clinical question of interest, and not for the choice of data collection and method of analysis to determine the estimand.” The estimand is defined first, from the clinical question. Only then is an aligned estimator chosen.
Then diagnose support, using the longitudinal checks rather than the point-treatment ones. Estimate the policy contrast last.
Run it the other way round — fit first, then ask what was estimated — and you get the case at the top of this lesson: a clean-looking number that nobody can say the meaning of. In regulated work you also get an analysis whose ordering an adopted guideline explicitly warns against.
- 1
Specify strategy
Static, dynamic, initiation, or sustained treatment.
- 2
Represent history
Time-varying covariates, treatment, censoring, outcome.
- 3
Choose g-method
Weight, simulate, or model treatment blips.
- 4
Diagnose support
Cumulative probabilities, weights, histories, and extrapolation.
- 5
Estimate policy contrast
Marginal outcome, curve, or structural parameter.
Comparison
Nothing here dominates, so the question is which model you would defend
No single g-method is better than the others in general, and looking for one is a way of avoiding the real question. What separates them is where each concentrates its modelling risk.
A marginal structural model puts the weight of the argument on the treatment and censoring models. Get those wrong and the pseudo-population is wrong, however tidy the final fit looks. Cole and Hernán's ladder is that risk made visible: six defensible answers from one dataset, separated only by a choice about weights.
G-computation moves the risk onto the outcome and covariate models. It has to simulate the world forward, and it will happily simulate a world that could not exist. 78,746 women in the Nurses' Health Study, followed over 1982–2002, were the first large-scale epidemiologic application of the parametric g-formula, published in the International Journal of Epidemiology in 2009. The observed 20-year risk of coronary heart disease in that cohort was 3.50%. The abstract reports what the simulation gave instead: “Under a joint intervention of no smoking, increased exercise, improved diet, moderate alcohol consumption and reduced body mass index, the estimated risk was 1.89% (95% confidence interval: 1.46-2.41).” Hold the two figures apart. The 3.50% was measured on 78,746 real women. The 1.89% was read off a simulated population, generated by the fitted covariate and outcome models, under a five-part intervention nobody in the cohort was assigned to. Both are numbers. Only one of them is an observation.
G-estimation asks you to get the shape of the blip right as well as the treatment model.
So the ranking is not a property of the methods. It is a property of your system, and of which parts of it you can honestly claim to know.
MSM/IPTW
Models treatment and censoring processes.
- Marginal target
- Weight instability
- Outcome model optional
Parametric g-formula
Models longitudinal covariate and outcome evolution.
- Simulates strategies
- Many models
- Flexible policy outputs
G-estimation
Models structural treatment effects and assignment residuals.
- Blip interpretation
- Complex implementation
- Useful for dynamic regimes
Steps
Run two of them on one problem and watch where each one flinches
Take a single dynamic treatment problem — one where the rule responds to the patient's evolving state rather than fixing the treatment in advance. There is a published one to work against. The HIV-CAUSAL Collaboration ran it in Annals of Internal Medicine in 2011, and states its design in one sentence: “Prospective observational data from the HIV-CAUSAL Collaboration and dynamic marginal structural models were used to compare cART initiation strategies for CD4 thresholds between 0.200 and 0.500 × 10(9) cells/L.” That is the shape of the exercise: a rule that fires when a measured state crosses a line, compared against other lines.
Start with the two sample sizes, because they are the exercise in miniature. Of 20,971 therapy-naive HIV-infected people, 8,392 entered the analysis. Everything downstream is a statement about the second number. The distance to the first is what requiring usable histories costs.
Then read the two outcomes against each other. Against the 0.500 threshold, the mortality hazard ratio was 1.01 (95% CI 0.84–1.22) at 0.350 and 1.20 (0.97–1.48) at 0.200. For AIDS-defining illness or death it was 1.38 (1.23–1.56) at 0.350 and 1.90 (1.67–2.15) at 0.200. Same people, same strategies, same weights. On mortality the 0.350 threshold is indistinguishable from 0.500. On the composite outcome it plainly is not. The choice of outcome moved the conclusion further than the choice of method will.
Now implement a second method on that same question. Hold the strategy, the covariate history and the horizon identical, so that the only thing changing is where the modelling sits. Record, for each one, the point at which a model is carrying the answer: the treatment and censoring models for the weighted route, the simulated covariate paths for the g-formula route. Those are the places where the estimate is a claim about the model rather than a reading of the data. Seeing both versions side by side is the fastest way to find them.
- 1
Define regime
History-dependent action and follow-up horizon.
- 2
Fit treatment models
For MSM weights and sequential support.
- 3
Fit evolution models
For g-computation of covariates and outcomes.
- 4
Run diagnostics
Weights, simulated histories, calibration, and influence.
- 5
Compare policy outcomes
Point estimates, intervals, and sensitivity.
Example
Four terms that get used as if they were interchangeable
The four have been circling each other since the case. They are easy to blur, because they answer the same question and share the same assumptions. What distinguishes them is what each one actually models. Each also has a published analysis behind it you can go and read.
- A marginal structural model states the contrast you want — the outcome under one treatment history against the outcome under another — and hands the confounding problem to the weights rather than solving it inside the model. The marginal structural Cox model of the 2000 Epidemiology paper is the founding applied instance. It returned 0.7 where the baseline-adjusted analysis had returned 2.3.
- A treatment-history weight is a single number per person, built by multiplying the inverse probabilities of each successive decision. That is precisely why it can grow enormous without any one step looking wrong: 46.8, in a set of stabilised weights averaging 1.04.
- G-computation is less an estimator applied to the data than a simulation of the data, run forward under an intervention you specify, with the outcome read off the simulated population. Hence 1.89% under a five-part intervention, against 3.50% observed, in 78,746 women of the Nurses' Health Study.
- G-estimation recovers the parameters of a structural nested model from the part of treatment assignment that the recorded history did not predict. Treatment residuals are its raw material, not a diagnostic. PCP prophylaxis and survival, in the observational sub-study embedded in AIDS Clinical Trials Group Trial 002, is where Robins and colleagues did it, in Epidemiology in 1992. They reported the width of what a structural nested failure time model could settle: “We find that, under our assumptions, the data are consistent with prophylaxis therapy increasing survival by 16% or decreasing survival by 18% at the alpha = 0.05 level.”
How to choose, and what agreement between methods is actually worth
Reach for a marginal structural model when you can credibly estimate how treatment and censoring were determined at each step, and when the quantity you owe someone is a marginal comparison between strategies. Reach for g-computation when the outcome and the evolution of covariates are the parts you can model with a straight face. Consider a structural nested model when the effect is naturally expressed as a blip at each moment, or when the regimes you care about are dynamic.
The HIV-CAUSAL question was in fact answered twice, on one dataset, by two g-methods. That makes it the right place to test what agreement is worth. The same 8,392 individuals went through the parametric g-formula instead of IP weighting, in a 2011 analysis by Young and colleagues. The 5-year mortality risks came out at 2.65% (95% CI 2.15–3.43) for the CD4 500 threshold, 3.06% (2.59–3.67) for 350 and 3.65% (3.01–4.44) for 200. The ordering matches the weighted analysis and the intervals are narrower. They say plainly what bought the narrowing: “However, estimators based on the parametric g-formula are more efficient than IP weighted estimators. This is often at the expense of more parametric assumptions.”
So the tempting move at the end — run two methods and treat their agreement as confirmation — is encouraging, and it is not proof. Those two analyses rest on the same 8,392 histories, the same sequential exchangeability assumption and the same gaps in support. The 12,579 people who did not enter the analysis are missing from both of them equally. What the pair does settle is where each method spends its risk. The weighted route bought wider intervals and asked less of the outcome model. The g-formula route bought narrower ones and asked more. That is a real finding about the two estimators. It is not a second measurement of the world.
Pick the method whose weakest model is the one you would be willing to defend out loud.
Key takeaways
- A marginal structural model gets its answer from weights built out of whole treatment histories, not out of single decisions. The same MACS data gave 3.6 crude, 2.3 baseline-adjusted and 0.7 under a marginal structural Cox model.
- The g-formula works from the other end. It models how the data evolve and simulates the population forward under a strategy — a 1.89% simulated 20-year coronary heart disease risk against 3.50% observed, in 78,746 Nurses' Health Study women.
- G-estimation targets the parameters of a structural nested model, using what is left over when treatment is predicted from the past. On ACTG Trial 002 that delivered a range from an 18% decrease to a 16% increase in survival at α = 0.05.
- All of these methods need sequential exchangeability and positivity, and positivity is model- and parameter-specific. The four repairs — smaller adjustment set, different projection function, restricted sample, modified intervention — each change the question rather than rescue it.
- Support can collapse across a sequence of decisions even when every individual decision looks perfectly common. Stabilised weights averaging 1.04 in 918 people still reached 46.8, and truncating them moved the estimate 0.32 units while the standard error fell by a third.
- A comparison between methods is only worth running if it shows where each one leans on a model instead of hiding it. The IP-weighted and g-formula analyses of the same 8,392 HIV-CAUSAL individuals agree partly because they share every assumption and every missing history.