Skip to content
AI.info

Causal inference

Association, Prediction, Intervention, and Counterfactuals

Separate observed association, future prediction, intervention effects, and unit-level counterfactual reasoning.

By the end you can

Example

The advertising that returned 4,173% until someone switched it off

In March 2012 eBay stopped buying paid search advertising on its brand keywords — the queries that contain the word “ebay” — on Yahoo! and MSN. The equivalent campaign kept running on Google, as a control. Then 99.5 percent of the paid click traffic eBay had stopped paying for came straight back through unpaid natural search. The company had been buying clicks from people who were already on their way to eBay.

A second test in the same study was larger. It ran on non-brand keywords, with the ads switched off in roughly 30 percent of US Nielsen DMAs. Conventional OLS on the pre-experiment data implied a return on investment of 4,173% without time and geographic controls, and 1,632% with them. The experimental estimate was −63%, with a 95% confidence interval of [−124%, −3%]. It was computed against eBay's imputed US paid-search spend of $51 million a year.

Same firm, same channel, same period, same spend. The number was +4,173% when an analyst conditioned on who had clicked. It was −63% when the firm intervened and decided who saw an ad. The three economists who ran the experiment put the mechanism in their abstract: “Because search clicks and purchase behavior are correlated, we show that returns from paid search are a fraction of conventional non-experimental estimates. As an extreme case, we show that brand-keyword ads have no measurable short-term benefits.”

Nobody had to click. The people who did were already shopping, already loyal, already headed for the site. Every one of those things raises purchasing on its own, ad or no ad. The gap in the logs was real. It belonged partly to the advertising and partly to the kind of person who reaches for a link like that. No amount of staring at the recorded data separates the two.

What eBay had before the experiment was one number that four different questions could lay claim to. Only one of those questions was the one the spend was meant to answer.

  • The association is plain enough. In the recorded data, people who clicked the paid ads bought far more than people who did not, and the pre-experiment regression turned that gap into 4,173% — 1,632% once time and geographic controls were added.
  • As prediction it holds up too. Knowing that someone clicked a paid ad for “ebay shoes” genuinely helps forecast a purchase, and it keeps helping even after the ad turns out to have caused none of it.
  • An intervention means switching the ads off, or on, by a rule the firm controls. That is the geo-randomised study eBay eventually ran in roughly 30 percent of US Nielsen DMAs. It returned −63%, with a 95% confidence interval of [−124%, −3%].
  • The counterfactual is what those same clickers would have bought had the ad never been shown to them. No log anywhere contains it. The brand-keyword shutdown recovered 99.5 percent of the forgone traffic in aggregate, which is a fact about the population, not about any one shopper.

Analogy

Watching a queue versus changing its rules

Priority customers wait less. The timestamps say so every day, and nobody disputes it.

Forecasting tomorrow's average wait from those same timestamps is a second thing you can do with them, and it works well enough to staff the counter by.

Changing who gets into the priority lane is a third thing, and this is where the timestamps run out. Some of the people in that lane were quick to serve in the first place, which is part of why they were put there. Move a slow customer into it and the lane is no longer the lane you measured. The wait grows for everyone standing behind. Reading eBay's click logs and then setting the advertising budget was exactly this: reading a clock, then reaching for the rules.

Watching a system and rewriting its rules do not produce the same numbers, even when the same clock records both.

Reading a column is not the same act as setting it

Filtering is what you do to data you already have. Take the rows where the ad was clicked, take the rows where it was not, compare them. Nothing in the world moved while you did that. You sorted records that were already written. Whatever decided who ended up in which pile is still deciding.

Setting is different. You choose who gets the treatment, by a rule you control, and in doing so you break the old process. Intent no longer selects the treated group, because you selected it. That is precisely what eBay did when it turned the ads off in a random set of markets.

The distinction has a name and a notation. Causal questions sort into three levels — Association, Intervention, Counterfactual — and each level gets its own symbol: P(y|x) for Association, P(y|do(x),z) for Intervention, P(y_x|x′,y′) for Counterfactuals. The typical counterfactual question is “What if I had acted differently?” Judea Pearl set the scheme out in Communications of the ACM in 2019. One sentence carries the ordering: “The second level, Intervention, ranks higher than Association because it involves not just seeing what is but changing what we see.”

Four questions hide inside most claims, and it helps to say them in order. What co-occurred in the data already recorded? What can be forecast for a case nobody has seen yet? What changes when a policy decides who gets treated? And what would have differed for the same people under a different strategy? The first two draw on the same recorded rows and sit at the first level. The third is the do-operator. The fourth is the last rung.

That the rungs do not quietly collapse into one another is a proved theorem, not a preference. The formal Causal Hierarchy Theorem measures the models where they do collapse, with respect to the Lebesgue measure over a suitable encoding of L3-equivalence classes of structural causal models, and finds that “the subset in which any PCH collapse occurs is measure zero”. Four researchers at Columbia's Causal AI Lab published it. Corollary 1 states the consequence in one line: “To answer questions at Layer i, one needs knowledge at Layer i or higher.”

So the first two questions you can answer with the data on hand. The last two you cannot, unless the treatment was assigned in a way you are willing to defend out loud.

P(y|x) and P(y|do(x)) are different quantities, and by Corollary 1 no quantity of the first ever adds up to the second.

Comparison

One sentence, four claims, four amounts of evidence

Hormone therapy protects the heart. Four people can say that sentence and mean four incompatible things by it. This particular sentence has had its readings pulled apart in public, on one cohort, with both sets of numbers printed.

The observational Nurses' Health Study appeared to show that women who took hormones were protected from coronary heart disease. Hernán and colleagues re-analysed that same cohort as a sequence of emulated trials, matching the protocol of the randomised Women's Health Initiative. The protection went away: “The ITT hazard ratios (HRs) (95% confidence intervals) of CHD for initiators versus noninitiators were 1.42 (0.92-2.20) for the first 2 years, and 0.96 (0.78-1.18) for the entire follow-up.” Nothing new was collected. The same women, the same records, a different question — and an answer close to the randomised WHI result of 1.29 (1.02–1.63) for coronary heart disease.

The analyst means the pattern recorded among the nurses. Read as a description of who took what, it is right. The forecaster means that hormone-use history carries signal about later coronary events. Also right, and it needs nothing further to become right.

The clinician means that prescribing hormones will lower a woman's risk. That claim leans entirely on how hormone use came to be initiated. When the initiation was re-cast as a trial protocol, the estimate moved to 1.42 (0.92-2.20) in the first two years and 0.96 (0.78-1.18) overall.

The fourth reader means something narrower and much heavier: that a particular woman who had a heart attack would not have had it, had she been prescribed hormones. Neither the cohort nor the trial speaks to that. The other version of that woman was never recorded.

FigureComparison · 3 columns

Association

Recorded coaching status co-occurs with activity.

  • No action implied
  • Sensitive to selection
  • Useful for description

Prediction

Coaching status helps forecast later activity.

  • Validated on future cases
  • May exploit confounding
  • Does not isolate effect

Intervention

Assigning coaching changes mean activity.

  • Requires causal design
  • Defines policy and timing
  • Supports a decision

Example

The same drug, said four ways

Advertising has no monopoly on this. The same four questions sit inside a sentence about medicine, and there the gap between the first reading and the third was large enough to stop a trial early.

The Women's Health Initiative estrogen-plus-progestin trial randomised 16,608 postmenopausal women aged 50–79, 8,506 to hormones and 8,102 to placebo. It was halted on 31 May 2002, after a mean 5.2 years. The hazard ratios came back the wrong way. With nominal 95% confidence intervals: 1.29 (1.02–1.63) for coronary heart disease, 1.26 (1.00–1.59) for invasive breast cancer, 1.41 (1.07–1.85) for stroke, 2.13 (1.39–3.25) for pulmonary embolism. Decades of observational evidence had pointed the other way. The conclusion in JAMA is one line long: “Overall health risks exceeded benefits from use of combined estrogen plus progestin for an average 5.2-year follow-up among healthy postmenopausal US women.”

The words that separate the four readings of a drug sentence are small enough to slide past a reader in a hurry. A tense here, a verb there, and the evidence required has quietly multiplied.

  • Women who took hormones had lower recorded rates of coronary heart disease. That describes who was taking what and explains nothing about why.
  • Hormone-use history improves a risk forecast, and it keeps doing so even if no one can say what the mechanism is.
  • Assigning the therapy would change average risk. That is a claim about what follows when a clinician acts, and the claim the 8,506 hormone assignments were built to test. The answer was 1.29 (1.02–1.63) for coronary heart disease.
  • This patient's stroke would not have happened under placebo. That asserts something about a life that was not lived. The trial's 1.41 (1.07–1.85) is a contrast between two arms, not a difference measured in any one woman.

Steps

Split a claim into the four questions it might be hiding

The exercise takes a minute. It is worth running on any claim that shows up wearing a verb like improves, drives, identifies, or prevents.

Write the sentence down. Then write it four more times: once as a statement about data already collected, once as a forecast about cases not yet seen, once as a claim about what happens when someone acts, and once as a claim about what would have happened to the same people under a different action. Change nothing but the question being asked.

Take a real one. The claim is that a widely used risk algorithm identifies the patients who need extra care. The algorithm scored need by predicting future health-care cost. Obermeyer and colleagues examined it across millions of US patients. As a forecast of spending it worked. As a statement about need it did not: at any given score, Black patients were substantially sicker than White patients. Correcting the target would have raised the share of Black patients flagged for extra help from 17.7% to 46.5%. Two New York State regulators wrote jointly to the CEO of UnitedHealth Group about Optum's Impact Pro on 25 October 2019: “By relying on historic spending to triage and diagnose current patients, your algorithm appears to inherently prioritize white patients who have had greater access to healthcare than black patients.” The letter demanded the company prove the algorithm was not discriminatory or stop using it.

That is the exercise done in public, by a regulator. The sentence held at the second level and broke at the third. Three of the four versions will usually turn out to be things your evidence does not reach. That is the useful part of the exercise, not a failure of it. The version you can defend is the version to say out loud, and it is often the dullest one on the page.

FigureProcess · 5 steps
  1. 1

    Record the observation

    State the measured association without causal verbs.

  2. 2

    Define the forecast

    Name the future target and deployment population.

  3. 3

    Specify the intervention

    Describe who sets what, when, and how.

  4. 4

    State the counterfactual

    Name the alternative strategy and outcome horizon.

  5. 5

    Choose the evidence

    Match design and validation to the intended claim.

Key idea

Counterfactual language can create false certainty

Only one outcome ever happens to a person. The patient took the drug or did not. The shopper saw the ad or did not. Whichever it was, the other version went unrecorded and always will.

A model will hand you an individual number anyway, and it will be a number about the world as it was administered. One rule-based learning system was trained on a pneumonia dataset of 14,199 patients — 46 features, 9,847 for training and 4,352 for test. The rule it wrote was “HasAsthma(x) ⇒ LowerRisk(x)”. Asthmatic pneumonia patients were usually admitted not just to hospital but directly to the ICU, so their recorded mortality came out below average. Caruana and colleagues describe the trap in 2015: “The bad news is that because the prognosis for these patients is better than average, models trained on the data incorrectly learn that asthma lowers risk, when in fact asthmatics have much higher risk (if not hospitalized).”

The pattern was true. As a per-patient prediction under the historical policy it was accurate. As a guide to who could safely be sent home it was lethal, because sending the patient home is a different policy from the one that generated every row of the training data.

Accuracy did not save it either. The most accurate models in that study were neural nets — AUC 0.86 against 0.77 for logistic regression on one of the pneumonia datasets, and about 0.02 better on this one. The team decided not to field them, “not because the asthma problem could not be solved, but because the lack of intelligibility made it difficult to know what other problems might also need fixing.” On the same data the intelligible GA2M model from that study reaches AUC 0.8576, against 0.8432 for logistic regression. It rediscovers the same asthma term.

So describe a per-person score as what it is. Say which population the estimate is aimed at, how uncertain it is, whether anyone comparable appeared in the data at all, and what has to be true about how treatment was assigned before the number means what it appears to mean.

A per-person effect is a conditional average under the policy that produced the data, not a difference anybody was ever in a position to measure.

Check the verbs before you fit anything

The check that catches all of this is a language check, and it costs a few seconds. Words like because, effect, identifies, prevented, and would have are promises. Each commits you to naming an intervention and naming the comparison it is measured against. If you cannot name both, the sentence is not ready. The repair is to say something smaller: what was observed, or what can be forecast.

In drug regulation this is no longer advice. A trial has to write its estimand down in advance. The rule is ICH E9(R1), the addendum on estimands and sensitivity analysis in clinical trials, adopted on 20 November 2019 and in effect in the EU since 30 July 2020. Its glossary defines the estimand: “A precise description of the treatment effect reflecting the clinical question posed by the trial objective. It summarises at a population-level what the outcomes would be in the same patients under different treatment conditions being compared.” With it goes a declared strategy for every intercurrent event — treatment policy, composite, while on treatment, hypothetical, or principal stratum. The verbs have to be settled before the data are analysed, in two jurisdictions, in writing.

Done habitually, the check guards two ends at once. A prediction that works stops being sold as a treatment, which is the mistake eBay's pre-experiment arithmetic walked into and the mistake the pneumonia rule encoded. And an average that holds across a population stops being retold as the story of one particular person, which is the mistake waiting at the far end of the same lesson.

No model catches either one. The sentence does.

A causal verb is a claim about a study you either ran or did not run — and since 20 November 2019 the regulator wants it written down first.

Key takeaways