Causal inference
Causal Inference Capstone: Design and Defend an Intervention Study
Integrate estimands, DAGs, target trials, experiments, observational methods, quasi-experiments, causal ML, policy, equity, and reporting.
By the end you can
- Translate a real decision into a complete causal estimand and target-trial protocol
- Choose and defend a design from randomized, observational, quasi-experimental, or combined evidence
- Build diagnostics, sensitivity, transport, equity, and policy-value plans
- Write a decision memo with explicit conditions for action or non-identification
The final product is an evidence system, not a treatment-effect number
You will want to hand in a number. The number is the last thing a reviewer looks at, and it is the easiest part of the work.
Look at what a defensible one costs. Remote monitoring after a heart-failure hospitalization has already been through a randomized trial. BEAT-HF, published in JAMA Internal Medicine in 2016, reports an adjusted hazard ratio of 1.03 for 180-day all-cause readmission, 95% CI, 0.88-1.20, P = .74. That line took 1,437 patients hospitalized for heart failure. They were randomized at six California academic medical centers between 12 October 2011 and 30 September 2013. Monitoring of weight, blood pressure, heart rate and symptoms plus nurse health-coaching calls went to 715 of them; usual care to 722. The estimate is one sentence of the abstract. Everything that entitles anyone to read it is the other years.
What you actually hand in is a chain. At one end sits a decision somebody has to make. At the other sits a program that runs, gets watched, and one day gets switched off. In between are the versions of the intervention you actually mean, the quantity that would answer the decision, how people were assigned to it, and how the outcome was measured. Then why this data can identify that quantity at all, how you estimated it, and what the diagnostics said. Then what policy will and will not permit, and who keeps the program honest afterwards. Any one of those links can cap what you are allowed to claim. How careful the rest of it was does not matter.
Which is why honest scope is rewarded here and precision is not. A capstone that ends in redesign the study, we can only bound this, randomize the next round, or do not deploy is stronger work than a tight interval resting on an assumption nobody defended.
Expertise shows up here as friction: the unjustified action gets harder to take, and the supported one still goes through.
Analogy
A flight-readiness review rather than an engine test
A plane does not leave the gate because the engine passed its test. Somebody signs off on the structure, the sensors, the weather limits, the crew, the abort procedure and the maintenance record. Any one of those can hold the aircraft on the ground while the engine sits there working perfectly.
Your estimator is the engine.
One item on that checklist has no equivalent on a runway, and the difference is not a small one. An inspector can put a hand on a bolt. Nobody can inspect a counterfactual. You will never see the patient who was treated standing next to the same patient, the same week, untreated. BEAT-HF got as close to that as anyone gets: 715 patients assigned to monitoring and 722 to usual care, by chance rather than by judgement. Even then the two groups are only comparable in the aggregate, never patient by patient. That is why everything else on the checklist has to carry more weight here than it does on a runway.
So the question a capstone answers is never whether the estimate is good. It is whether anyone reading the file can see what would have grounded the flight.
Release is a property of the whole system, and the estimate only gets a vote in it.
Example
The hospital's monitoring program has already been through a randomized trial: BEAT-HF
A hospital wants to send recently discharged patients home with remote monitoring, hoping to keep them out of the emergency department for the next thirty days. The idea is a good one. The first question is not how to analyse it. It is whether somebody has already run it.
Somebody has. BEAT-HF randomized 1,437 patients hospitalized for heart failure between 12 October 2011 and 30 September 2013, at six California academic medical centers. Monitoring of weight, blood pressure, heart rate and symptoms plus nurse health-coaching calls went to 715 of them; usual care to 722. The result section of the abstract says: “The intervention and usual care groups did not differ significantly in readmissions for any cause 180 days after discharge, which occurred in 50.8% (363 of 715) and 49.2% (355 of 722) of patients, respectively (adjusted hazard ratio, 1.03; 95% CI, 0.88-1.20; P = .74).” There was no significant difference in 30-day readmission either, or in 180-day mortality.
Now read the observational analysis the hospital was about to commission against that. Monitoring has been handed out before, and never at random: the sickest patients and the best-connected ones got it. Monitoring also changes how often a nurse picks up the phone. In BEAT-HF the health-coaching calls were part of the intervention, not a nuisance to be adjusted away. Patients wear the thing diligently at first and less as the weeks pass. And when a patient does turn up at an emergency department across town, the hospital may never hear about it. The outcome goes quietly missing for exactly the people it matters most for. On top of that, there are only enough devices for one-third of the patients who qualify.
Suppose you defeated every one of those threats. You would arrive, expensively, at a quantity that a randomized comparison of 715 patients against 722 has already put at 50.8% versus 49.2%. The estimator was never the binding constraint. What is left to decide is which version of monitoring you mean, which patients you are talking about, whether you can randomize or must emulate a trial, and how the outcome gets measured. Then which patients have comparable counterparts at all, whether the effect differs between them, who gets the scarce devices, and whether that allocation is fair. Then what you watch after launch, and what would make you stop.
- The decision on the table is not whether monitoring works in general. For one written-down version of it — the device readings plus nurse health-coaching calls — BEAT-HF answered that at 50.8% (363 of 715) against 49.2% (355 of 722). The live decision is who gets a device, given there are only enough for one-third of the eligible patients.
- The quantity that answers that decision is the thirty-day effect on emergency visits of one explicitly written monitoring strategy, with harms counted on the same page as benefits. Not the effect of monitoring, which is a word covering at least the device, the calls, and the nurse who reads the readings.
- One design will not carry it. Run a pragmatic experiment where randomizing is possible, an observational analysis emulating that experiment to reach patients it could not, and a longer follow-up to see whether the effect survives. BEAT-HF's own primary window closed at 180 days, not at thirty.
- Consent, who may see the data, what happens if a subgroup is harmed, how a patient appeals, and when the program is retired belong inside the study design, not in a memo written after launch.
Visual
Five stages, and a reviewer who reads them backwards
The dossier has five stages, and they run in the order the work runs: the question and protocol, the causal structure, the design portfolio, the estimation and diagnostics, then the policy and lifecycle.
A reviewer reads them the other way. They start at the recommendation, ask which decision it changes, and walk back down the chain until they reach a link that will not take their weight. Walk BEAT-HF backwards and every step holds. The conclusion rests on 1.03 (95% CI, 0.88-1.20). That rests on a randomized comparison of 715 patients against 722. That rests on a named intervention and a stated 180-day window, in a population defined by hospitalization for heart failure at six named-in-advance sites.
Build the map so it can be walked in that direction. Write each assumption where it is actually used, and mark each decision point. Do that and you find the weak link before the reviewer does. That is the only version of this that ends well.
Question and protocol
Decision, population, strategies, timing, outcome, estimands.
Causal structure
DAG, variable roles, interference, measurement, selection.
Design portfolio
Randomized, observational, quasi-experimental, and transport evidence.
Estimation and diagnostics
Nuisance models, support, sensitivity, falsification, uncertainty.
Policy and lifecycle
Capacity, equity, rollout, monitoring, redress, stop and retirement.
Example
Four documents that argue for you when you are not in the room
Four terms get confused with each other constantly, and they do different jobs.
- A decision memo says what you recommend, on what evidence, with how much uncertainty, what the alternatives were, and who owns the call.
- An assumption register is the versioned list of what your identification rests on and what supports each item, so that when one assumption moves you can see which claims moved with it.
- A design portfolio is the set of experiments and observational or quasi-experimental analyses you are running together, each with a stated job.
- A stop rule is written before launch and names the condition that would make you pause the program, redesign it, or withdraw the causal claim. Here is what its absence buys. The Epic Sepsis Model was deployed at hundreds of US hospitals. It was validated from outside in 2021, across 38,455 hospitalizations of 27,697 patients at Michigan Medicine. Its area under the ROC curve was 0.63 (95% CI, 0.62-0.64). The results section of that abstract reads: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” The program ran at scale with nobody holding a condition that would have switched it off.
Key idea
The capstone can fail by collecting every method instead of building one coherent design
The most common way this work fails is by being impressive. The team fits propensity scores to the historical monitoring data, grows a forest to hunt for heterogeneous effects, adds a DiD around the month the program expanded, appends a page of sensitivity metrics, and submits something thick.
Read closely, those pieces answer different questions about different populations under assumptions that cannot all hold at once. Nothing in the report says which one the recommendation actually rests on. And none of them, stacked, outweighs one randomized comparison of 715 patients against 722. Thickness is not evidence, and it is not even a proxy for it.
Pick one primary design and say so out loud. Everything else earns its place by doing a named job: testing a specific threat to that design, carrying its result to a population it did not cover, or telling the policy owner something they need in order to decide. A method with no job is decoration, and decoration is the first thing an adversarial reader removes.
If you cannot say in one line why a method is in the dossier, it is not evidence, it is volume.
Example
Four artifacts, each of which has to survive being read alone
Each one should be short enough that somebody reads it to the end, and complete enough that somebody else can replay what you did from it. Two of the four have published specifications behind them, which is worth knowing before somebody dismisses them as house style.
The first is one sentence long. Hernán and Robins wrote it in 2016: “Causal inference from large observational databases (big data) can be viewed as an attempt to emulate a randomized experiment—the target experiment or target trial—that would answer the question of interest.” Specify the trial you would run in full protocol detail. Then judge your observational analysis by how well it emulates that trial.
- An estimand and protocol card, naming the trial you would run if you could, the versions of the strategy you actually mean, and the clock that starts and stops the outcome. This is regulated ground. ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, was adopted under Step 4 on 20 November 2019. The EMA's CHMP adopted it on 30 January 2020, and it came into effect on 30 July 2020. Its glossary defines an estimand as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective. It summarises at a population-level what the outcomes would be in the same patients under different treatment conditions being compared.” It separates the estimand, the estimator and the estimate. And it requires the estimand to be built from five attributes: treatment, population, variable or endpoint, handling of intercurrent events, and a population-level summary.
- An assumption map: the DAG, where you have support and where you do not, how people were selected in, whether one patient's treatment can affect another patient's outcome, how the outcome was measured, and what would have to hold to carry the result somewhere else.
- An analysis plan naming the primary estimator before results exist, the alternatives you will also run, the diagnostics, the sensitivity checks, and how you handle testing many things at once.
- A decision memo covering what the policy is worth, the capacity you actually have, who benefits and who is left out, how it rolls out, what gets monitored, and what stops it.
Comparison
Four endings, all of which pass
The capstone is not graded on whether the hospital deploys anything. It is graded on whether the ending matches the evidence that was gathered.
Recommend the program as written when three things hold: the quantity you estimated is the one the decision needs, the design identifies it, and the value survives an honest push on the assumptions. Recommend it narrowed when it holds only for part of the population. Deploy where the support exists, and say plainly where it does not. Recommend more evidence first when the weak link is fixable. That usually means randomizing the next expansion, measuring the outcome you have been losing, or reporting bounds instead of a point. Recommend against deploying when the chain breaks somewhere you cannot repair — or when the chain holds perfectly and the answer is simply no.
That fourth ending is in print. Telephone-based interactive-voice-response telemonitoring was tested in Tele-HF, which randomized 1,653 recently hospitalized heart-failure patients: 826 to telemonitoring, 827 to usual care. The primary end point was readmission for any reason or death from any cause within 180 days. It occurred in 52.3% versus 51.5%: a difference of 0.8 percentage points, 95% CI, -4.0 to 5.6, P=0.75. Readmission for any reason was 49.3% versus 47.4%, and death 11.1% versus 11.4%. The authors closed their abstract in the New England Journal of Medicine in 2010 with the ending itself: “Among patients recently hospitalized for heart failure, telemonitoring did not improve outcomes. The results indicate the importance of a thorough, independent evaluation of disease-management strategies before their adoption.”
Only the last of those feels like failure. It is the one the hospital most needs to hear. It is also the hardest to write, because it means having done the work and still saying no.
Deploy progressively
Evidence supports a bounded policy.
- Randomized validation
- Guardrails and monitoring
- Rollback ready
Redesign study or intervention
Target or implementation is not yet coherent.
- Clarify versions
- Improve measurement
- Collect missing comparisons
Partial identification/no claim
Evidence supports only bounds or uncertainty.
- Honest limitation
- Protects decisions
- Guides next study
Steps
Hand it to a team that wants it to be wrong
Give a second team one instruction: attack the link you are proudest of.
Their job is to tell the story in which your effect is not real, using nothing but your own dossier. The sickest patients were the ones who got devices. The nurse calls did the work and the device was along for the ride. The patients who abandoned the device are the ones who ended up in an emergency department you never saw.
Notice what randomizing does to those three. It closes the first outright. It turns the second into a question about what your intervention is defined to include, rather than a bias to be adjusted away — BEAT-HF simply put the calls inside the intervention and said so. It does nothing at all to the third. An emergency visit that never reaches your records is missing from the outcome whether you randomized or not. If the second team can tell that story and your evidence cannot rule it out, that is your result, and you found it in time.
- 1
Attack the intervention
Versions, adherence, timing, and operational feasibility.
- 2
Attack identification
Hidden causes, selection, support, interference, and measurement.
- 3
Attack estimation
Nuisance error, influence, multiplicity, and unstable heterogeneity.
- 4
Attack transport and policy
Target differences, capacity, equity, and implementation.
- 5
Define response
Revise, bound, collect evidence, pause, or reject the claim.
Key idea
The questions that can stop a release
A review board earns its keep by asking short questions the team cannot answer quickly. Which version of monitoring are you claiming works? What makes the comparison fair? Are there patients in your target group with no comparable counterpart anywhere in the data? Would a missed emergency visit ever reach your outcome? Does any of this carry to a hospital that is not this one? Who is worse off the day it launches? And can the wards actually run it on a busy night?
Two of those have measured answers, and both are worse than teams expect.
Start with the missed emergency visit. Somebody counted them. Between July 2008 and September 2009, across California, Florida and Nebraska, 5,032,254 index hospitalizations among 4,028,555 patients ended in discharge. Within 30 days, 17.9% (95% CI, 17.9%-18.0%) of those hospitalizations were followed by at least one acute care encounter. Treat-and-release ED visits ran at 97.5 per 1,000 discharges (95% CI, 97.2-97.8). Readmissions ran at 147.6 per 1,000 (95% CI, 147.3-147.9). Seeing that took a linkage of state inpatient and emergency-department databases across the three states, reported in JAMA in 2013. The conclusions section of the abstract puts it this way: “After discharge from acute care hospitals in 3 states, ED visits within 30 days were common among adults and accounted for 39.8% of postdischarge hospital-based acute care visits. Improving care transitions should focus not only on decreasing readmissions but also on ED visits.” That share — 39.8% (95% CI, 39.7%-39.9%) of 1,233,402 encounters — is what an inpatient-readmission outcome never counts.
Now who is worse off at launch. A commercial risk algorithm was selecting patients for extra care-management resources, which is the same job as handing out a scarce device. Obermeyer and colleagues audited it in Science in 2019. Because it predicted health care costs rather than illness, Black patients at a given risk score were considerably sicker than White patients. Their abstract states the size of it: “Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%.” The rule was passing its own metric the entire time.
The board has done its job when it names the weakest link, not when it blesses the method the team liked best. A review that ends in agreement reviewed nothing.
You are not defending a method in that room. You are defending a recommendation, and it has to still stand once somebody has tried to break it.
Steps
Assemble it for readers who will never read it together
Submit one package, built so that each of its readers can work through it without you in the room and without needing each other.
A scientific reviewer should be able to find the estimand, the assumptions and the diagnostics, and reproduce what you did from them. The standard to aim at is a BEAT-HF-style report, where the intervention, the arms, the sites, the enrollment window and the endpoint can all be read off the first page. A policy owner should be able to find the recommendation, what it costs, who it reaches and what would stop it, without wading through the estimator. A patient, or somebody speaking for them, should be able to find out what was done with their data, what the program does to someone like them, and how to object to it.
The Tele-HF authors asked for thorough, independent evaluation of disease-management strategies before their adoption. Independent evaluation is only possible against a package somebody outside the team can actually read. A section that only makes sense to the person who wrote it is not finished.
1. Protocol and estimands
Target trial, intervention versions, population, outcomes, horizons, and effect scales.
2. Assumption and data map
DAG, measurement, support, selection, interference, and transport evidence.
3. Analysis and diagnostics
Primary estimator, alternatives, uncertainty, falsification, and sensitivity.
4. Decision and lifecycle
Policy value, equity, rollout, monitoring, redress, stop rules, and retirement.
The bar for recommending action, and what to do when you cannot clear it
Recommend action when all of this holds at once: the quantity you estimated is the one the decision needs, the design credibly identifies it, comparable patients exist across the range you would deploy to, the measurement would survive being checked, and the value still looks worth it after the assumptions are pushed as far as a critic plausibly would push them.
All of it. Not most of it.
That last push has a unit, and refusing to report it is a choice. VanderWeele and Ding introduced the E-value in 2017 for exactly this: “The E-value is defined as the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the treatment and the outcome to fully explain away a specific treatment-outcome association, conditional on the measured covariates.” They propose that in all observational studies intended to produce evidence for causality the E-value be reported, or some other sensitivity analysis be used. They suggest calculating it twice: once for the observed association estimate, after adjustment for measured confounders, and once for the limit of the confidence interval closest to the null. One number for the estimate, one for the edge of the interval. That is the whole cost of showing how hard your finding is to break.
When one of the conditions fails you still have something to hand in, and it is not a hedge. Narrow the target to where the evidence actually lives. Ask for the stronger experiment. Ask for the measurement you were missing. Report bounds rather than a point. Or state that there is no causal claim to be made here yet.
Then write down the part that is almost always left out: who owns the next decision, and what specific evidence would change this one. A recommendation without those can never be wrong, which is another way of saying it can never be checked.
Readiness is measured by how easy you made it for the next person to find your flaw.
Key takeaways
- A causal study is a chain running from the definition of the intervention to the decision it changes, and it is only as strong as its weakest link. BEAT-HF's 50.8% versus 49.2% is believable because 1,437 patients were randomized at six California academic medical centers between 12 October 2011 and 30 September 2013.
- The estimator is one part of the work, sitting alongside the design, the measurement, the diagnostics and whoever governs the program afterwards. The Epic Sepsis Model was deployed at hundreds of US hospitals, then validated at an area under the ROC curve of 0.63 (95% CI, 0.62-0.64) across 38,455 hospitalizations at Michigan Medicine.
- Randomized, observational, quasi-experimental and causal-ML methods are not competitors. They solve different problems, and the target trial of Hernán and Robins is the device that holds an observational analysis to the protocol a randomized one would have had.
- Sensitivity, falsification, support and transport belong in the main review, not in an appendix nobody reaches. VanderWeele and Ding propose that in all observational studies intended to produce evidence for causality the E-value be reported or some other sensitivity analysis be used.
- What a policy is worth includes the harms it causes, the capacity you actually have, who gets left out, and what it costs the people who run it. Correcting one commercial risk algorithm's target variable would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%.
- Redesigning, reporting bounds, and refusing to deploy are conclusions an expert is allowed to reach. Tele-HF's authors reached the third across 1,653 randomized patients, at 52.3% versus 51.5% on the primary end point, and asked for independent evaluation before adoption.