Kinds of learning
Prediction, Causation, and Counterfactual Questions
Distinguish predictive tasks from causal and counterfactual questions, with attention to confounding, policy effects, experiments, and decision relevance.
By the end you can
- Distinguish prediction from causal effect estimation
- Explain confounding and why conditioning can introduce bias
- Identify when randomized or quasi-experimental evidence is needed
- Recognize counterfactual claims that ordinary supervised models cannot justify
Key idea
An algorithm reaching roughly 200 million people predicted cost and was read as need
Being accurate about the process that generated the data does not make a model the right thing to act on. One commercial risk-prediction algorithm decides which patients are enrolled in extra care. By industry estimates it is applied to roughly 200 million people in the United States each year. It does not predict who is sickest. It predicts future health-care cost, and cost was treated as a stand-in for need. Obermeyer and colleagues took the system apart in Science on 25 October 2019, and their abstract names the mechanism: “The bias arises because the algorithm predicts health care costs rather than illness, but unequal access to care means that we spend less money caring for Black patients than for White patients.”
The consequence is measured, not asserted. Patients are auto-identified for enrolment at the 97th percentile of risk score. At that threshold Black patients carried 26.3% more chronic illnesses than White patients — 4.8 distinct conditions against 3.8, P < 0.001. Changing the label would raise the share of Black patients receiving extra help from 17.7% to 46.5%. The manufacturer replicated the finding on its own national dataset of 3,695,943 commercially insured patients. On the day of publication the New York State Department of Financial Services and Department of Health wrote jointly to UnitedHealth Group about Optum's Impact Pro.
Prediction asks what is likely under the observed process. A causal question asks something else. It asks how outcomes would change under an intervention.
The algorithm was doing what it was fitted to do — predicting spending — and the deployment read that number as how sick you are.
Case
The pneumonia model that learned asthma patients were lower risk
A mislabelled target is one route to that failure. Learning the record of past treatment is another. In the 1990s a rule-learning system was trained on a pneumonia dataset, and one of the rules it produced was “HasAsthama(x) ⇒ LowerRisk(x)” — pneumonia patients with a history of asthma were at lower risk of dying than the general population. Caruana and colleagues recount the episode in a 2015 paper on intelligible models for health care. The pattern was real, and its cause was the hospital. Such patients “usually were admitted not only to the hospital but directly to the ICU”, and the aggressive care was effective enough to reverse their prognosis. A model built to decide who could safely be sent home would have read the record of that care as a reason to withhold it.
Comparison
Three questions that use similar data but need different evidence
The wording of the question determines which assumptions and evaluation designs are relevant. Optum's Impact Pro answered the first of the three well enough to be deployed at national scale. It was used as though it had answered the second.
Predictive
What outcome is likely for this case under the current process?
- Uses associations for forecasting
- Evaluated on future-like observations
- Can exploit stable proxies
- Example: who will be readmitted
Causal
How would an intervention change an outcome on average or for a subgroup?
- Requires treatment and comparison logic
- Confounding threatens observational estimates
- Experiments can strengthen identification
- Example: effect of a reminder
Counterfactual
What would have happened to this case under a different action?
- Concerns an unobserved alternative world
- Requires stronger assumptions
- Individual claims are especially difficult
- Example: outcome without the treatment
Confounding creates predictive association without intervention value
A variable can influence both the action and the outcome. Sicker patients may receive more treatment and also have worse outcomes, which makes treatment appear harmful in raw observational data.
A predictive model can use that association effectively without resolving the causal story. Causal estimation cannot. It requires assumptions about which variables block or create bias, and those assumptions have a price that has been measured on data where the right answer was already known.
The National Supported Work Demonstration was randomized, so its answer was known. In 1982 dollars the experimental estimates were $851 (SE $317) of extra 1979 earnings for AFDC female participants and $886 (SE $476) for male participants. Then LaLonde threw the randomized control group away. Writing in the American Economic Review in September 1986, he rebuilt the evaluation with the observational comparison groups and the econometric adjustments then standard, and scored the results against those known benchmarks. For the women the non-experimental estimates were usually positive and larger than the experimental one. For the men they were negative and smaller. Same data, same programme, and the error running in opposite directions by subgroup. His abstract draws the conclusion: “This comparison shows that many of the econometric procedures do not replicate the experimentally determined results, and it suggests that researchers should be aware of the potential for specification errors in other nonexperimental evaluations.”
Case
16,608 women randomised, and an association with the wrong sign
Medicine has one very expensive worked example. Observational studies had linked postmenopausal hormone therapy to lower coronary risk, and it was prescribed on that basis for years. Then the Women’s Health Initiative randomised 16,608 postmenopausal women to estrogen plus progestin or placebo. On 31 May 2002, after a mean 5.2 years of follow-up, the data and safety monitoring board recommended stopping the trial. The results appeared in JAMA on 17 July 2002: hazard ratios of 1.29 (95% CI 1.02–1.63) for coronary heart disease, 1.41 (1.07–1.85) for stroke and 1.26 (1.00–1.59) for invasive breast cancer, against 0.66 (0.45–0.98) for hip fracture and 0.63 (0.43–0.92) for colorectal cancer. The association had been large, stable and repeatedly reproduced. On the outcome that mattered most it also had the wrong sign.
Visual
Ways to strengthen a causal claim
Every method leaves one requirement in place. Treatment, outcome, population, timing, and interference still have to be defined.
External variation is the rung that is easy to describe and hard to find. Card and Krueger found some. On 1 April 1992 New Jersey raised its minimum wage from $4.25 to $5.05 an hour, and Pennsylvania's stayed at $4.25. They surveyed 410 fast-food restaurants on both sides of that line, before and after the change, and reported it in 1993: “Relative to stores in Pennsylvania, fast food restaurants in New Jersey increased employment by 13 percent.” The standard model predicted the opposite. In 2021 the Royal Swedish Academy of Sciences awarded David Card a share of the Nobel Prize in Economic Sciences for this style of natural-experiment evidence.
The randomized rung is bounded by ethics and operations, and occasionally those bounds hand you the design. Oregon reopened its OHP Standard Medicaid programme in 2008 with about 10,000 slots and roughly 140,000 eligible adults. The state's own final report on what it did next gives the reason: “When the department decided to open enrollment for OHP Standard, an estimated 140,000 Oregon adults would have been eligible for the program, so the department determined the most equitable way to select enrollees was to create a reservation list from which names would be randomly drawn.” More than 90,000 people put their names on the list. Eight staggered draws from March to October 2008 sent applications to 35,476 people, and 10,031 enrolled.
Because assignment was random, a rationing decision produced a genuine randomized trial. After about two years it found no statistically significant effect of Medicaid on measured blood pressure, cholesterol or glycated hemoglobin. It cut depression by 30 percent, 9.2 percentage points from a base of 30, and it virtually eliminated catastrophic out-of-pocket spending, down 4.5 points from a base of 5.5. Randomization did not make every hoped-for effect appear. It made the answer legible in both directions.
Domain theory
Specify a causal graph or substantive mechanism that makes assumptions visible.
Observational adjustment
Control for measured confounders using a defensible design.
Natural or quasi-experiment
Use external variation that approximates randomized assignment.
Randomized experiment
Assign interventions by chance within ethical and operational limits.
Replication and sensitivity
Test alternative assumptions, populations, and interference patterns.
Example
Prediction and intervention in product design
One product can need both predictive and causal models, and the second is a separate piece of work with its own design and its own bill. The tutoring version of the question below was settled by assignment rather than by forecasting. Twelve authors reported two randomized controlled trials of Saga Education's high-dosage tutoring in Chicago public high schools in the American Economic Review in March 2023: 2,633 ninth- and tenth-graders randomized in the first, 2,710 students in the replication. Their abstract: “Participating in math tutoring increases math test scores by 0.18 to 0.40 standard deviations and increases math and non-math course grades.” An earlier working paper, from March 2021, reported the same two trials as 0.16 SD and 0.37 SD, at a cost of $3,500 to $4,300 per participant per year. No improvement in the accuracy of the first bullet in each pair below would have produced those numbers.
- Prediction: estimate which users are likely to cancel a subscription.
- Causal question: estimate who would remain because of a retention offer, which means withholding it from a comparable group.
- Prediction: forecast which students will fail a course.
- Causal question: estimate whether tutoring changes outcomes for comparable students — answered in Chicago by randomizing 2,633 students and then 2,710 more.
- Prediction: rank patients by expected readmission risk.
- Causal question: estimate the effect of a follow-up program on readmission.
Steps
Turn an intervention claim into a research design
Before modeling, make the hypothetical comparison explicit.
Step 5 is not a matter of taste, and at least one regulator has written the ordering into public documents. The 21st Century Cures Act, enacted on 13 December 2016, defines real world evidence as data on a drug's usage, benefits or risks “derived from sources other than randomized clinical trials”, and it required a framework within two years. The FDA published that framework in December 2018. It records that randomization has been considered a critical element in establishing a causal relationship, and then says what the alternative costs: “Although observational studies may provide credible evidence, there is a stronger scientific justification for deriving evidence of a drug effect from randomized controlled trials as compared to observational studies.” The statute defines a whole category of evidence by what it is not: not a randomized clinical trial. The framework then prices the difference.
1. Define the action
Specify who can receive what intervention, at which time, and with what version.
2. Define the outcome
Choose a measurement window and avoid outcomes influenced by later selection.
3. Name the counterfactual
State the alternative action or policy being compared.
4. Draw assumptions
Identify common causes, mediators, colliders, and interference.
5. Choose evidence
Use randomization, quasi-experiment, or justified observational adjustment.
6. Test sensitivity
Report how conclusions change under plausible unmeasured bias or policy variation.
Analogy
Umbrellas and rain
People carrying umbrellas are more likely to encounter rain. Umbrellas predict rain because forecasts and clouds influence the decision to carry one; forcing everyone to carry an umbrella would not cause rainfall.
Predictive association has the same separation from intervention. Umbrellas have one obvious confounder; real causal systems can carry many, along with feedback and interference among people.
Predictive power does not identify the effect of changing the predictor.
Causal language should match the design
Words such as “drives,” “prevents,” “because,” and “impact” imply more than predictive association. Feature importance, correlation and model explanations do not establish intervention effects.
How often the wording outruns the design has been counted, not merely deplored. Haber and his co-authors, in the American Journal of Epidemiology in 2022, screened 1,170 non-randomized articles from 18 high-profile journals published 2010–2019 and had three reviewers rate each abstract. The exposure–outcome linking language carried no causal implication in 13.8% of abstracts, weak in 34.2%, moderate in 33.2% and strong in 18.7%. The commonest linking word was “associate”, at 45.7%. The recommendations ran ahead of the findings they rested on: “The implied causality of action recommendations was higher than the implied causality of linking sentences for 44.5% or commensurate for 40.3% of articles.” At larger scale, Wang and Yu classified 91,933 observational-study abstracts in PLOS ONE on 12 August 2026 and found 31.7% making causal claims. That is close to the 31% Cofield and colleagues had found by hand in 525 obesity and nutrition papers.
Use descriptive language when the evidence is predictive. If the product acts on causal claims, involve domain experts and appropriate experimental or causal-inference methods.
Case
eBay switched its brand ads off: a 4,100% return measured as −63%
eBay priced the difference between the two vocabularies. In March 2012 it switched off its paid search advertising on brand keywords across Yahoo! and MSN and measured what was lost: “almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search.” The experiment reached Econometrica in 2015. Across the wider set of experiments it compared the two ways of valuing the advertising channel. Standard non-experimental attribution produced “a ROI of over 4,100% without time and geographic controls, and a ROI of over 1,400% with such controls”. The experimental estimate was “a ROI of −63%, with a 95% confidence interval of [−124%, −3%], rejecting the hypothesis that the channel yields positive returns at all.” Every click in the observational estimate was genuine. The ads were being credited with traffic the company already had.
Figure
Position
An attribution report is not a measurement of what the spending did
Attribution dashboards, feature-importance charts and channel reports share one habit: they credit an outcome to a lever by comparing the people the lever touched with the people it did not. eBay's paid search advertising is the cleanest demonstration of what that comparison is worth. Non-experimental attribution valued the channel at over 4,100% return, or over 1,400% once time and geographic controls were added. Switching the advertising off measured −63%, with a 95% confidence interval running from −124% to −3%. The gap is not bad data. It is that 99.5% of the forgone clicks were immediately picked up by natural search, from people who were arriving anyway.
The other cases in this lesson are the same error in other clothes. Asthma predicted lower pneumonia mortality because patients with asthma were sent straight to intensive care. Cost predicted need for roughly 200 million people a year, until someone checked the illness counts at the 97th percentile and found 4.8 conditions against 3.8. Hormone therapy predicted lower coronary risk across large, stable, repeatedly reproduced observational studies, and the trial that randomised 16,608 women reported a hazard ratio of 1.29 in the other direction. And when LaLonde rebuilt a training-programme evaluation without its randomized control group, the standard econometric adjustments overshot the $851 experimental benchmark for women and turned negative for men.
One question separates evidence from arithmetic whenever a number is offered as what something is worth. What was compared with what, and was anything actually changed? If nothing was switched off, delayed, or randomly assigned, the report describes how the organisation has behaved. It does not say what the lever does.
The same channel came out at over 4,100% by attribution and −63% by experiment, and every click in both numbers was real.
Key takeaways
- Prediction estimates what is likely under an observed process, while causation concerns intervention effects.
- A strong predictor can be a proxy, consequence, or policy artifact rather than a cause: cost stood in for need at the 97th percentile, and asthma stood in for a bed in the ICU.
- Confounding links intervention assignment and outcomes through common causes, and adjustment can fail loudly — LaLonde's non-experimental estimates overshot the $851 experimental benchmark for women and went negative for men.
- Counterfactual questions require assumptions about unobserved alternative outcomes.
- Randomized and quasi-experimental designs can strengthen causal identification: a state border and a wage law for Card and Krueger, a rationing lottery of 35,476 applications for Oregon.
- Causal language should not exceed the evidence supplied by the research design, and in Haber's sample the action recommendations were more causal than the findings in 44.5% of articles.