Skip to content
AI.info

Causal inference

Causal Inference as Decision Science

Frame causal inference as a disciplined chain from a decision and intervention to an estimand, evidence, assumptions, and action.

By the end you can

Example

The same data said 316% and 73%, and only one of them was an effect

Fifteen US advertising experiments on Facebook had already done the expensive part. Randomisation decided which users were eligible to be shown the ad and which were not. Four researchers got all fifteen to work with: 500 million user-experiment observations, 1.6 billion ad impressions. They published what they found in Marketing Science in 2019.

The point was not to measure the ads. It was to ask what an analyst would have concluded from the very same records, if the experiment had not been there to consult. Take the one they label Study 4. The randomised contrast — the average effect of the treatment on the treated — put the lift at 73%. Now set the randomisation aside. Keep every row, compare the users who were exposed with the users who were not, and the lift reads 316%. Tighten that comparison the way a careful team would, by exact matching on age and gender, and it still reads 222%.

Of the 316% the authors write: “This estimate represents the combined lift due to treatment and selection and is more than four times the lift due to treatment of 73%.”

Study 4 was not picked for its drama. Fourteen of the studies had a checkout-conversion outcome. In 7 of those 14, the observational point estimate missed the experimental one by more than a factor of three.

Nothing in the observational analysis was broken. The rows were real, the arithmetic was right, and the comparison was the one anyone would have made. What it could not do was separate what the advertising did from the fact that the people who end up exposed are not a random slice of the audience. The two arrive fused, and no column in the table says in what proportion. The exposed-versus-unexposed number was an honest answer to a question nobody had asked.

  • The decision on the table was narrow and practical: run this campaign or do not. The number that answers it is the randomised 73%, not the 316% the same rows produce once the randomisation is ignored.
  • The intervention had to be one specific thing: this ad, served to this eligible audience, over this window. A label as broad as advertising has no single effect to be measured.
  • The comparison had to be those same eligible users left unexposed. The experiment builds that group by randomising. The observational table only resembles it, and exact matching on age and gender moved the estimate from 316% no further than 222%.
  • The failure was reading an association — “the combined lift due to treatment and selection” — as the result of an action nobody had taken.

Comparison

Three questions were sharing one dataset, and only one of them was the campaign question

Those same records could answer three quite different questions honestly. The three are checked in three different ways.

Describe what happened and you are reporting counts: this many users were exposed, that many converted. You check the claim by going back to the records. With 500 million user-experiment observations and 1.6 billion ad impressions behind it, a descriptive claim of this kind is about as solid as data gets. Predict who converts next and you are ranking people who have not decided yet. You check that by holding rows out, fitting on the rest, and seeing whether the ranking survives contact with users the model never saw. A model can pass that test fairly and completely.

Ask what changes if the campaign runs, and no held-out set can help. The comparison the question needs was never run. There is no hidden slice of the data containing the exposed users as they would have been unexposed. That is exactly what the 15 experiments manufactured, at cost. It is also why 316% and 73% can both be computed from one dataset while only one of them is an effect.

The gap between what teams confidently expect from a change and what randomisation reports is itself a measured quantity. Six Microsoft researchers stated it flatly in 2013: “Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve.” Kohavi and Thomke restated the split in Harvard Business Review in 2017 — one third effective, one third neutral, one third negative. They added that only about 10% to 20% of experiments at Google and Bing generate positive results.

Three questions, one table, three standards of proof. A result validated to the predictive standard and presented to the causal one looks identical on a slide. The 316% is what that looks like.

FigureComparison · 3 columns

Description

Summarizes patterns in observed data.

  • Target: what occurred
  • Checks: measurement and sampling
  • No intervention claim

Prediction

Forecasts an outcome for new cases.

  • Target: future accuracy
  • Checks: generalization
  • Association may be enough

Causal inference

Contrasts outcomes under interventions.

  • Target: counterfactual effect
  • Checks: design assumptions
  • Supports action only within scope

The contract has a name and seven components

The contract the advertising analyst never wrote down has a citable form. Hernán and Robins set it out in the American Journal of Epidemiology in 2016. Their abstract states the move directly: “Causal inference from large observational databases (big data) can be viewed as an attempt to emulate a randomized experiment—the target experiment or target trial—that would answer the question of interest.”

The target trial is not a metaphor. It is a protocol with seven components, every one of them written down before the analysis begins: eligibility criteria, treatment strategies, assignment procedures, follow-up period, outcome of interest, causal contrast(s) of interest, and analysis plan. Fill those seven in and the question has a shape a colleague can argue with. Leave them blank and there is nothing to argue with. That is not the same thing as agreement.

Regulators fix a similar list, with dates attached. ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, was adopted on 20 November 2019 and took legal effect in Europe on 30 July 2020. It pins the target quantity through five attributes: treatment condition, population, variable or endpoint, handling of intercurrent events, and population-level summary. It also says what such a quantity is. “An estimand is a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective. It summarises at a population level what the outcomes would be in the same patients under different treatment conditions being compared.”

That named quantity has to exist before the method is chosen, because the data can never simply hand it over. Every user is either exposed or not; the other version of them is recorded nowhere. Bridging that gap is a separate job with a separate name. Identification asks whether the design and the assumptions you are willing to defend pin the target down at all. Estimation asks, once it is pinned down, how precisely a finite sample can measure it. Exact matching on age and gender is estimation performed carefully on a target that was never identified. It returned 222% where the experiment returned 73%.

The written components do the work, not the size or richness of the dataset. That has been shown on one of the most argued-over datasets in epidemiology. Hernán, Robins and six co-authors re-analysed the observational Nurses' Health Study as a sequence of emulated trials, in Epidemiology in 2008. Written down that way, the same records returned intention-to-treat hazard ratios for coronary heart disease of 1.42 (95% CI 0.92–2.20) in the first 2 years and 0.96 (0.78–1.18) over the whole follow-up. That is close to the randomised Women's Health Initiative result rather than a contradiction of it. Their conclusion: “Our findings suggest that the discrepancies between the Women's Health Initiative and Nurses' Health Study ITT estimates could be largely explained by differences in the distribution of time since menopause and length of follow-up.”

Nothing was added to the data between the two readings. Eligibility, time zero and the follow-up window were specified differently, and the answer moved.

The same Nurses' Health Study records gave 1.42 in the first 2 years and 0.96 over the whole follow-up; the eligibility, time-zero and follow-up rules chose between them, not the data.

Example

Every one of these is the campaign question in different clothes

The shape recurs everywhere a team can already predict who is at risk and now has to decide whether to act. In each case the prediction is available and the answer is not. In the first two, the answer arrived only when someone paid for the randomised comparison.

  • Hormone therapy was prescribed for years on observational evidence suggesting benefit. Then it was randomised. The Women's Health Initiative gave estrogen plus progestin to 8,506 postmenopausal women aged 50–79 with an intact uterus, and placebo to another 8,102 — 16,608 in all. Coronary heart disease came back at a hazard ratio of 1.29 (nominal 95% CI 1.02–1.63; adjusted 0.85–1.97), invasive breast cancer at 1.26 (nominal 1.00–1.59; adjusted 0.83–1.92). In absolute terms: 7 more CHD events and 8 more invasive breast cancers per 10,000 person-years. On 31 May 2002, after a mean 5.2 years of follow-up, the data and safety monitoring board recommended stopping. The NHLBI halted the trial, and the results were released on 9 July 2002, in JAMA. The conclusion: “The risk-benefit profile found in this trial is not consistent with the requirements for a viable intervention for primary prevention of chronic diseases, and the results indicate that this regimen should not be initiated or continued for primary prevention of CHD.”
  • Whether health coverage improves health was settled for one population by a lottery. Oregon ran a Medicaid lottery in 2008. The Oregon Health Insurance Experiment compared 6,387 adults randomly selected to apply with 5,842 not selected. After roughly two years there was no significant effect on blood pressure, cholesterol or glycated hemoglobin, while depression screening fell by 9.15 percentage points (95% CI −16.70 to −1.60; P=0.02). The report, in the New England Journal of Medicine in 2013: “This randomized, controlled study showed that Medicaid coverage generated no significant improvements in measured physical health outcomes in the first 2 years, but it did increase use of health care services, raise rates of diabetes detection and management, lower rates of depression, and reduce financial strain.”
  • A marketplace considering a change to seller fees needs to know whether completed transactions would rise, and whether supply would quietly drain away while they did. Two outcomes, one change, and no reason to assume the sign is the same for both. The one-third split reported at Microsoft is the base rate for how often such a change does what its authors expect.
  • An operations team weighing preventive maintenance against the schedule it already runs needs the failures the change would avoid. That is not the same list as the failures it can foresee.

Visual

The chain starts at a decision and has to end at one

Follow the campaign question through and you get five links. The first and the last are the same kind of thing.

It starts with a decision someone is actually going to make: run the campaign or do not. That decision fixes an estimand. The estimand link is nothing more elaborate than writing out “an attempt to emulate a randomized experiment—the target experiment or target trial—that would answer the question of interest.” The seven target-trial components and the five ICH E9(R1) attributes exist so that this link cannot be left vague: population, treatment condition, comparison, outcome, horizon, how intercurrent events are handled, how the result is summarised at population level.

Identification comes next, and asks the uncomfortable question. Given how exposure actually came about, can this contrast be recovered from this data, under assumptions we are willing to state and defend? In the Facebook comparison this is exactly where the chain broke. Exposed and unexposed users differ before any ad is served. So what the table offers is “the combined lift due to treatment and selection” rather than the lift due to treatment. No amount of care further down repairs that.

Only if the answer is yes does estimation begin. It produces a number and, just as importantly, an honest interval around it. Note the order. 222% is a more carefully estimated version of the wrong quantity, not a partial version of the right one.

Then the chain returns to where it started, because the number is not the deliverable — the decision is. Only one third of tested ideas “improved the metric(s) they were designed to improve,” so carrying the estimate back to the run-or-not choice is where most of the value sits. A chain that breaks at the identification link should stop there loudly, rather than proceeding quietly to a well-estimated answer to some other question.

FigureProcess · 5 steps
  1. 1

    Decision

    Name the action and the choice it is meant to inform.

  2. 2

    Estimand

    Define the population-level counterfactual contrast.

  3. 3

    Identification

    State why observed data can reveal that contrast.

  4. 4

    Estimation

    Choose an estimator and quantify uncertainty.

  5. 5

    Decision

    Combine the estimate with harms, costs, capacity, and values.

Steps

Write the protocol before anyone opens the data

Take one question your own team is currently arguing about — a product change, a policy, an offer someone wants to send — and write its protocol on a single page, before any modelling.

Use the seven components Hernán and Robins list, in their order, and do not leave one blank because it seems obvious: eligibility criteria, treatment strategies, assignment procedures, follow-up period, outcome of interest, causal contrast(s) of interest, and analysis plan. The treatment strategy has to be specific enough that an operations colleague could execute it without a follow-up question. The assignment procedure is where you write down honestly how people actually come to be treated in your setting. That is the clause the 316% quietly skipped.

Then check the same page against the five attributes ICH E9(R1) fixes: treatment condition, population, variable or endpoint, handling of intercurrent events, and population-level summary. The fourth is the one teams forget. It asks what you will do about the people who stop, switch, or receive something else halfway through. A page that survives both lists is “a precise description of the treatment effect reflecting the clinical question posed by a given clinical trial objective,” with your question in place of the clinical one.

Then write the part most briefs omit. List the assumptions that would have to hold for your data to identify that contrast, and next to each one write what you would expect to see if it were false. The advertising case tells you the size of the stakes. When exposure is not as good as random within the variables you hold, the association overshoots, and 316% against 73% is how far the overshoot can run.

Finish with the sentence you would say to the decision-maker if those assumptions turn out to be indefensible. If that sentence is hard to write, you have found the real problem early. That is the cheapest place to find it.

FigureProcess · 5 steps
  1. 1

    Name the unit

    State who or what receives the intervention.

  2. 2

    Define strategies

    Specify treatment and comparison precisely.

  3. 3

    Set time zero

    Align eligibility, assignment, and follow-up.

  4. 4

    Choose the outcome

    Define measurement, horizon, and competing events.

  5. 5

    Declare the decision

    State what action the estimate could change.

Key idea

A precise estimate can answer the wrong causal question

Suppose the team goes away, gets a proper design, and comes back with a tight interval on the effect of the treatment. They may still have nothing usable. A label can cover several different actions at once — a deeper version and a shallower one, delivered at different moments, through different channels, to people at different points in their history. Each may do something different to the outcome. Average them and you get a number that describes no action anyone can actually take.

The clearest illustration is a label that was defensible and still turned out to be too broad. The Women's Health Initiative tested one thing precisely: estrogen plus progestin, in postmenopausal women aged 50–79 with an intact uterus, 8,506 of them against 8,102 on placebo. Its conclusion was correspondingly narrow — that regimen “should not be initiated or continued for primary prevention of CHD.” Not hormones in general, not every regimen, not every purpose. The observational cohort was later re-analysed as a sequence of emulated trials, and the long-standing disagreement between the two bodies of evidence turned out to sit inside clauses of exactly this kind: “differences in the distribution of time since menopause and length of follow-up.” Two literatures wearing the same label had not been estimating the same quantity.

This is why ICH E9(R1) carries an attribute that looks bureaucratic until you need it: handling of intercurrent events. People stop, switch, or start something else. A label that does not say what happens to them is not yet an estimand. The machinery will happily estimate the average of a bundle nobody defined, and it will report a small standard error while doing it. That is what the economist Charles F. Manski means by predictions that are fragile, “with conclusions resting on critical unsupported assumptions or leaps of logic.”

So write the decision and the estimand first, and choose the method afterwards. If two people reading the label would go out and do different things, the intervention is not yet defined operationally. Then either narrow the claim to the version you can define, or redesign the study around the version you care about.

ICH E9(R1) fixes five attributes for a reason: a treatment label that does not say what is done, to whom, and how disruptions are handled has nothing left to be an estimate of.

What a defensible causal conclusion sounds like

A conclusion worth acting on says out loud what it is about: the intervention as it would actually be delivered, the comparison it was measured against, the population it applies to, the horizon at which the outcome was read, the assumptions holding the whole thing up, and the uncertainty around the number.

The Oregon Health Insurance Experiment shows what that sounds like when the answer refuses to be a single thing. It compared 6,387 adults randomly selected in Oregon's 2008 Medicaid lottery with 5,842 who were not, and reported: “This randomized, controlled study showed that Medicaid coverage generated no significant improvements in measured physical health outcomes in the first 2 years, but it did increase use of health care services, raise rates of diabetes detection and management, lower rates of depression, and reduce financial strain.” Read that sentence for what it does rather than what it found. It keeps the outcomes separate instead of collapsing them into a verdict. And it stamps the horizon onto the null result — “no significant improvements in measured physical health outcomes in the first 2 years” — so that nobody can quote it as a claim about ten.

It then does one more thing, which is easy to forget. It separates the size of an effect from the worth of it. Depression screening fell by 9.15 percentage points while blood pressure, cholesterol and glycated hemoglobin did not move. What that combination is worth is a judgement for the people who have to fund it, not a statistical result. Hand over an estimate without that line and everyone will assume the analysis already made the call.

And sometimes the honest conclusion is that the target was never identified. That is a result, not a failure to produce one. Manski made the case in The Economic Journal in 2011: “Yet policy predictions often are fragile, with conclusions resting on critical unsupported assumptions or leaps of logic. Then the certitude of policy analysis is not credible.” His paper is a typology of how that certitude gets manufactured: conventional certitude, duelling certitudes, conflating science and advocacy, wishful extrapolation, illogical certitude and media overreach. His earlier NBER working paper, in 2010, had named only the first four. What he presses for instead is reporting what the data actually identify. A range rather than a point. A claim about the narrower group where the assumptions do hold. A proposal for the experiment that would settle it, or the plain statement that the effect is currently unknown. In a room about to act on 316%, that last sentence would have been worth more than the confident answer everyone left with.

An effect that is real and an effect that is worth buying are two separate findings, and only the first one comes out of the analysis.

Key takeaways