Causal inference
Time-Varying Treatments and Treatment–Confounder Feedback
Define longitudinal treatment strategies, time-varying confounding, treatment–confounder feedback, and sequential exchangeability.
By the end you can
- Represent treatment and covariate histories over discrete decision times
- Explain treatment–confounder feedback and why standard adjustment can fail
- State sequential consistency, exchangeability, and positivity
- Define static and dynamic longitudinal treatment strategies
Example
CD4 count was both the reason for the zidovudine and the result of it
A clinician watches a CD4 lymphocyte count drift the wrong way and treats. That is ordinary care. It is also where the trouble starts. The CD4 count that prompted the treatment had itself been pushed around by the zidovudine given before it.
The loop was measured in 2000, in the Multicenter AIDS Cohort Study. Compare treated men with untreated and the crude mortality rate ratio for zidovudine was 3.6 (95% CI 3.0-4.3). The drug looked lethal, because it went to the men who were already failing. Adjust for baseline CD4 and the other baseline covariates in the usual way and the ratio fell to 2.3 (1.9-2.8). Better, and still on the wrong side of one. A marginal structural Cox model that also controlled time-dependent confounding gave 0.7 (conservative 95% CI 0.6-1.0). Hernán and colleagues reported all three, in Epidemiology.
The arithmetic in every one of those analyses was fine. What differed was what each of them did with a single column of data. Their abstract says it in one sentence: “In this study, CD4 lymphocyte count is both a time-dependent confounder of the causal effect of zidovudine on survival and is affected by past zidovudine treatment.”
That is the double role. CD4 count was the reason later zidovudine was given, so leaving it out left the later doses confounded — the 3.6. It was also a consequence of earlier zidovudine. Put it into an ordinary regression alongside treatment and it soaks up the very improvement the earlier doses had produced. Adjust for it and you erase part of the effect. Omit it and you confound the rest. One regression cannot do both. The distance from 3.6 to 0.7 is the size of the mistake, and it is a change of sign, not a rounding. A measurement that steers the next decision while carrying the last one's effect is called treatment-confounder feedback. Everything below is about it.
- The history at any decision point is everything on record up to that moment: the treatments already given, the covariates measured, the outcomes so far, and the fact that the patient was still being observed at all. In the Multicenter AIDS Cohort Study, that means the zidovudine already taken and the CD4 counts already returned.
- A time-varying confounder moves over follow-up and predicts both what treatment comes next and how the patient ends up. That is exactly why it looks like something you should control for. CD4 count predicted who was treated next and who died. Controlling for it at baseline alone moved 3.6 to 2.3.
- Feedback is the extra turn of the screw. Earlier treatment changed that confounder, so the thing you would adjust for is partly the treatment's own handiwork. Only a method built for that, the marginal structural Cox model, returned 0.7 (conservative 95% CI 0.6-1.0).
- A dynamic strategy is not a fixed dose but a rule — start when CD4 first falls below a stated threshold, hold when it does not. What any patient receives depends on the history they have accumulated by then.
You are not estimating a treatment, you are estimating a course
Ask what the medication does and you have to finish the sentence. Does what, given how, for how long, and adjusted in response to what? When treatment can change at every visit, there is no single thing called the treatment for an effect to belong to. What a patient could have lived through is a whole course: this dose, then that one, then a pause. Potential outcomes here are indexed by entire sequences. More usefully, they are indexed by rules that say what to do given what has been seen so far.
At each decision point the clinician is reading covariates that have moved since the last visit, and those same covariates predict the outcome. That is what makes them look like confounders worth adjusting for. In a one-shot study they would be.
What changes is that they were also affected by prior treatment. Condition on them and you remove the part of the effect that travelled through them. You can also manufacture bias where none existed. The longitudinal g-methods were built for this exact shape of problem.
The first large-scale epidemiologic application of one of them shows what the shape buys you. Five things changed at once, over twenty years, each feeding back into the reasons for the others: no smoking, increased exercise, improved diet, moderate alcohol consumption, reduced body mass index. Run the Nurses' Health Study over 1982-2002 as things actually went and the 20-year risk of coronary heart disease was 3.50%. Under that joint intervention, estimated with the parametric g-formula, it was 1.89% (95% CI 1.46-2.41). Taubman and colleagues published the estimate in 2009.
The authors say plainly why the machinery is needed: “The parametric g-formula, which appropriately adjusts for time-varying confounders affected by prior exposures, is especially well suited to estimating effects when the intervention involves multiple factors (joint interventions) or when the intervention involves decisions that depend on the value of evolving time-dependent factors (dynamic interventions).” Their own example of such a decision is “start exercising if diagnosed with diabetes”. That is the treat-when-the-biomarker-worsens rule this lesson turns on. The g-methods pay for their answers with assumptions, and those assumptions must hold at every decision point, not once at the start.
One column of data can be the fix for later bias and the carrier of earlier effect at the same time, and nothing in a regression can ask it which role it is playing.
Example
These words collapse into each other if you let them
The loop is easy to lose hold of, and it is usually the vocabulary that slips first. A rule gets called a treatment. An assumption gets called a fact. That first distinction is not a methodologist's preference. It is written regulation. ICH E9(R1), the “Addendum on Estimands and Sensitivity Analysis in Clinical Trials”, was adopted on 20 November 2019 and came into effect in Europe on 30 July 2020. It defines an estimand as “a precise description of the treatment effect reflecting the clinical question posed by the trial objective”. Say what you are estimating before you estimate it. Four distinctions are worth keeping apart while reading on.
- Treatment history means the actual sequence of actions a patient received through follow-up, visit by visit, rather than a summary such as how much they got in total — the zidovudine record month by month, not an ever-treated flag.
- A dynamic strategy is a rule that reads the history and then decides. The same rule can produce quite different sequences in two patients whose measurements went differently. E9(R1) allows exactly this: the interventions compared “might be individual interventions, combinations of interventions administered concurrently, e.g. as add-on to standard of care, or might consist of an overall regimen involving a complex sequence of interventions.”
- Sequential exchangeability is the assumption that at every decision point, once you condition on the history recorded up to then, who got treated next was as good as arbitrary. It has to hold at each of those points, not on average across them.
- Treatment-confounder feedback names the situation from the opening case: prior treatment changed the very covariates that will justify the next decision, so the confounder is partly a product of the thing being studied. CD4 count in the Multicenter AIDS Cohort Study was both. The three estimates 3.6, 2.3 and 0.7 are what that cost.
Visual
The arrow everyone forgets points back into the covariate
Lay the pieces on a line and the difficulty shows itself. Baseline history sits at the left. Treatment at a given time follows from it. Updated covariates come next, and two arrows leave them rather than one: a short one forward to the next treatment, and a long one onward to the outcome.
The arrow that gets left off the drawing is the one running into those covariates from the treatment already given.
A published analysis lets you point at that box rather than describe it. ACCORD's 10,251 participants were followed for a median 3.4 years, and a post-hoc study asked which A1C measurement predicted death. The answer was the average on treatment. It beat both the last-interval A1C and the first-year fall. The intensive strategy looked worse than the standard one only where average A1C exceeded 7%. Riddle and colleagues reported that in Diabetes Care in 2010.
On-treatment A1C is downstream of the assigned strategy and predictive of death at the same time. Both arrows leave one box. Read the trial through that box and the achieved biomarker tells a different story from the randomised comparison of the two target rules. The authors draw the line carefully: “These analyses implicate factors associated with persisting higher A1C levels, rather than low A1C per se, as likely contributors to the increased mortality risk associated with the intensive glycemic treatment strategy in ACCORD.”
Follow the earlier treatment through the updated covariates to the outcome and you are looking at effect — the part of the benefit that adjustment would erase. Follow the updated covariates forward to both the next treatment and the outcome and you are looking at confounding — the bias that ignoring them would leave in. Both paths run through the same box on the timeline. So no choice about that box is clean. That is why the assumptions have to be inspected decision by decision rather than all at once.
- 1
Baseline history
Eligibility and pre-treatment covariates.
- 2
Treatment at t
Action chosen from current history.
- 3
Updated covariates
Response, toxicity, behavior, or severity.
- 4
Next treatment
New action informed by updated covariates.
- 5
Outcome
Result under the full treatment history.
Comparison
A fixed course and a responsive rule are different questions
What you can estimate depends on what you are willing to state in advance. A static strategy names the sequence up front and asks what would have happened if everyone had followed it whatever their measurements did along the way. A dynamic strategy names a rule instead and lets each patient's history decide when treatment fires.
The static contrast is easier to define and often easier to identify. Its weakness is that it may describe a course of treatment no clinician would ever deliver. Real clinicians respond to what they see, which is what created the feedback in the first place.
The dynamic contrast matches practice, and it asks more of the data in return. The HIV-CAUSAL Collaboration ran exactly this comparison: rules of the form start therapy when CD4 first falls below X, estimated with dynamic marginal structural models in 8,392 of 20,971 eligible therapy-naive patients. Take the 0.500 x 10^9 cells/L threshold as the reference. The hazard ratio for AIDS-defining illness or death was 1.38 (95% CI 1.23-1.56) at the 0.350 threshold and 1.90 (1.67-2.15) at 0.200. Waiting cost something, and the cost grew the longer you waited. Run the same two contrasts without adjusting for time-varying confounding and they come back at only 1.22 (1.11-1.34) and 1.38 (1.25-1.53). Cain and colleagues, writing in the Annals of Internal Medicine in 2011, put the direction of that gap plainly: “If we had not adjusted for time-varying confounding, these estimates would have been closer to the null.”
The same paper shows the price of asking a dynamic question. Every action the rule prescribes has to actually appear in the records, among people whose history calls for it. But 47% and 82% of eligible people never crossed the 0.350 and 0.200 thresholds at all, and only 8,392 of the 20,971 contributed to the comparison. A rule that says to treat at a level of the biomarker where nobody was ever treated is a rule the data has no answer for, however carefully the estimator is written.
This is not a technical preference between two estimands. It is a choice about which policy you are advising on.
Static strategy
Assign the same treatment sequence to all eligible units.
- Simple definition
- May be unrealistic
- Can violate support
Dynamic strategy
Assign treatment according to measured history.
- Policy-like
- Requires rule definition
- Needs longitudinal support
Observed practice
Follow whatever treatment occurred.
- Descriptive
- Confounded over time
- Not a strategy contrast
Analogy
A thermostat reads the room it just heated
The heat comes on because the room is cold, and the room, having been heated, reads warmer at the next check. That reading is doing two jobs at once. It is the reason for the next adjustment, and it is the record of what the last adjustment achieved.
Now judge the heating by comparing settings against readings while holding the reading fixed. The heat will look useless. Every gain it produced has already been absorbed into the number you are controlling for.
The thermostat need not stay a metaphor, because the setting itself has been randomised. ACCORD assigned 10,251 people with type 2 diabetes and a median A1C of 8.1% to two treat-to-target strategies: intensive, targeting A1C below 6.0%, or standard, targeting 7.0-7.9%. What was randomised was not a dose but a number to steer toward — a rule, applied again and again as the readings came back. It was registered on ClinicalTrials.gov before it ran.
The intensive arm was stopped early. There had been 257 deaths there against 203 on standard therapy (hazard ratio 1.22, 95% CI 1.01-1.46, P=0.04). The primary cardiovascular composite, meanwhile, showed 352 against 371 events (HR 0.90, 0.78-1.04, P=0.16). The deaths, not the composite, ended it. The trial report, in the New England Journal of Medicine in 2008, puts it in one line: “The finding of higher mortality in the intensive-therapy group led to a discontinuation of intensive therapy after a mean of 3.5 years of follow-up.”
A thermostat's rule is at least written down, and ACCORD's was written down and registered. A clinician's ordinarily rests partly on unrecorded judgment and shifting preference. That is precisely why exchangeability at each decision point stays an assumption rather than something you can verify outside a trial. Written down or not, once decisions repeat, the measurement you would adjust for is partly a thing your own decisions made.
Repeated decisions do not merely respond to the world; they rewrite the measurements you will later use to judge them.
Example
The loop is not a medical curiosity
Medicine has no monopoly on this pattern. It turns up wherever a decision gets made more than once and the next decision reads a number the last one moved. That is most places where an adaptive policy meets a record of what happened. Outside medicine it has been measured and named.
- Doses are raised or lowered on the strength of biomarkers that earlier doses had already shifted, so the clinical case is the general case: zidovudine and CD4 count in the Multicenter AIDS Cohort Study moved a mortality rate ratio from 3.6 to 0.7 depending only on how the feedback was handled.
- A lender revises a credit limit after watching repayment behaviour that the previous limit helped to produce.
- Schools intensify support for pupils whose measured performance already reflects the tutoring those pupils were given earlier.
- A recommender changes what someone sees. That changes what they click. Those clicks decide what the recommender shows next. Chaney and colleagues named this “algorithmic confounding” and simulated it in 2018: communities of 100 users run for 1,000 time intervals, ten new items introduced per interval, each configuration repeated over ten random seeds. Training on data already shaped by recommendations amplified homogenisation of user behaviour without a corresponding gain in utility. Their abstract states the shape exactly: “These systems are often evaluated or trained with data from users already exposed to algorithmic recommendations; this creates a pernicious feedback loop.”
Steps
Draw the timeline before you fit anything
Pick a problem where the same decision gets made again and again — ideally one of your own rather than a textbook's — and lay it out in time before writing a line of model code.
Mark the decision points first. At each one, write down what was known: which actions had already been taken, which measurements had come back, what the person deciding could actually see. That list is the history, and every assumption you will later make is conditioned on it.
Then find the covariates that appear on both sides, the ones steering the next decision and predicting the outcome. For each of them ask the single question this lesson turns on. Could an earlier action have moved it? CD4 count answers yes, and so does on-treatment A1C. Every yes is a feedback arrow. Every feedback arrow marks a variable that ordinary adjustment will handle wrongly in one direction or the other — the 3.6 or the 2.3, never the 0.7.
Finish by naming the strategy you actually want to compare against, in words a practitioner would recognise as a rule they could follow. Start when CD4 first falls below 0.350 x 10^9 cells/L is such a rule. Then check the records for people who were in fact treated that way. Cain and colleagues had to, and found that 47% and 82% of eligible people never crossed the 0.350 and 0.200 thresholds at all.
- 1
Choose decision times
Visits, days, transactions, or policy reviews.
- 2
List histories
Baseline, prior treatment, evolving covariates, observation.
- 3
Mark feedback
Which covariates are changed by prior treatment?
- 4
Define strategies
Static sequences or history-dependent rules.
- 5
State assumptions
Sequential exchangeability, positivity, consistency, censoring.
When to shrink the question instead of the model
There are two ways to get this wrong, and they look like opposites. Control only the baseline covariates and you ignore every reason later treatment was given, as though the whole course had been set on the first day — that is the 2.3 in the Multicenter AIDS Cohort Study. Put all the time-varying covariates into one regression and you cut the earlier treatment off from its own consequences, conditioning on variables that treatment produced.
Neither is fixed by tuning. The way out is to represent the treatment process as it actually ran in time, then choose a g-method matched to the strategy you meant to ask about, static or dynamic.
The signal that you are in this territory is plain enough. Treatment changes repeatedly, and evolving covariates guide the later changes. Once both are true, a baseline contrast or an ever-treated comparison is not a simplification of the policy question. It hides the policy, and the number it returns describes no course of treatment anyone could follow.
Sometimes the honest move is to ask less, and there is a well-documented case of that paying off. The Nurses' Health Study and the Women's Health Initiative trial had famously disagreed about estrogen/progestin therapy and coronary heart disease. So the observational cohort was reworked into an emulated randomised trial, restricted to the narrow question of initiating therapy. Reporting in Epidemiology in 2008, Hernán and colleagues gave intention-to-treat hazard ratios of 1.42 (95% CI 0.92-2.20) in the first 2 years and 0.96 (0.78-1.18) over the whole follow-up. Adherence-adjusted by inverse probability weighting, the same two came out at 1.61 (0.97-2.66) and 0.98 (0.66-1.49). Those figures sat close to the trial, and unlike the earlier Nurses' Health Study analyses of sustained use. The authors could then say why: “Our findings suggest that the discrepancies between the Women's Health Initiative and Nurses' Health Study ITT estimates could be largely explained by differences in the distribution of time since menopause and length of follow-up.” The narrow question, answered well, reconciled a famous disagreement that the broad one had produced.
Where histories are poorly measured — visits missing, covariates recorded whenever somebody thought to record them — the effect of a sustained strategy may simply not be identified, and no estimator recovers what was never observed. A narrower question about starting treatment, answered well, is worth more than a sustained-strategy figure the data cannot support. What matters is saying which of the two you did.
The unit of analysis is a rule for deciding, which means you have to be able to state the rule before you can estimate anything at all.
Key takeaways
- When treatment can change over time, what you are estimating is a whole treatment history, or a rule for choosing one as you go. ICH E9(R1), adopted on 20 November 2019, allows the treatment attribute of an estimand to be an overall regimen involving a complex sequence of interventions.
- The covariates that justify later treatment decisions are frequently the ones earlier treatment already moved: CD4 lymphocyte count in the Multicenter AIDS Cohort Study, and on-treatment A1C in the ACCORD post-hoc analysis of 10,251 participants.
- Ordinary adjustment breaks under that feedback, either blocking the earlier effect or leaving the later decision confounded. Zidovudine's mortality rate ratio ran 3.6 (95% CI 3.0-4.3) crude, 2.3 (1.9-2.8) with baseline adjustment, and 0.7 (conservative 95% CI 0.6-1.0) under a marginal structural Cox model.
- Sequential exchangeability must hold at every decision point given the history up to then, not once at baseline. Where it is not adjusted for, estimates drift toward the null: the HIV-CAUSAL contrasts fell from 1.38 and 1.90 to 1.22 and 1.38.
- A dynamic strategy can only be estimated if the records contain people treated the way the rule prescribes. Cain and colleagues could use only 8,392 of 20,971 eligible therapy-naive patients, and 47% and 82% of eligible people never crossed the 0.350 and 0.200 x 10^9 cells/L thresholds at all.
- When histories are badly measured, a narrower point-treatment question is the safer thing to answer. The Nurses' Health Study initiation emulation returned 0.96 (0.78-1.18) over full follow-up, close to the Women's Health Initiative trial and unlike the earlier sustained-use analyses.