Causal inference
Distributional Effects, Equity, and Causal Reporting
Analyze effect distributions, quantiles, subgroup harms, equity objectives, causal report structure, reproducibility, and evidence governance.
By the end you can
- Distinguish average effects from distributional and subgroup consequences
- Connect causal estimands with equity and resource-allocation choices
- Report assumptions, diagnostics, sensitivity, and external validity coherently
- Create reproducible causal evidence and decision records
Example
A $3,477 gain for the younger children, a reversal for the adolescents
The Moving to Opportunity housing-voucher experiment raised earnings on average. Inside that average, the sign of the effect flips. Children who moved below age 13 earned $3,477 more per year in their mid-twenties — 31% more, against a control mean of $11,270. Children who moved after age 13 had, if anything, negative long-term impacts. Chetty and two colleagues reported the split in the American Economic Review in 2016. Age at move was a group that existed before anyone was assigned anything.
An average taken across all of those children would have been arithmetically correct. It would have carried the younger movers' gain while absorbing the adolescents' loss without a trace. The number is not false. It is answering a narrower question than the one on the table, and nobody reading a one-line recommendation could tell that from the recommendation.
What a release needs is enough structure that a reader can rebuild the claim and find its weakest link without asking the analyst. That is not house style. Clinical trials have had a rule for it since 2019. ICH E9(R1), the addendum on estimands and sensitivity analysis, requires a trial's target of estimation to be fixed in advance. Five named attributes do the fixing: treatment condition, population, variable, intercurrent-event strategy, and population-level summary. Its glossary defines an estimand as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective. It summarises at a population-level what the outcomes would be in the same patients under different treatment conditions being compared.”
An estimand card carries those five attributes. An assumption register lists each identification claim, the evidence behind it, the person who owns it, and how far the answer moves when it fails. A support report says who the data could not speak for: the excluded groups, the tails of the weights, the places where the estimate is extrapolation rather than measurement. A decision memo carries the recommendation, the alternatives that lost, the uncertainty that remains, the monitoring, and the conditions under which the policy stops. Four documents, one job. Let a stranger find the weak link.
- The average effect is the mean contrast between what happened to the population and what would have happened without the policy. Across the Moving to Opportunity children, that one number would have pooled a $3,477 annual gain with an adolescent loss.
- A distributional effect asks instead how the whole spread of outcomes moved, quantile by quantile, rather than where its center landed.
- Subgroup harm is what the children who moved after age 13 took. A real negative direction inside one slice of the population, reported as negative long-term impacts, that a mean absorbs without a trace.
- An equity objective is the rule you bring to that fact rather than read out of it. It says whose outcomes count and which losses you refuse to trade away — a choice ICH E9(R1) cannot make for you, even after the estimand is fixed.
Causal evidence informs values; it does not choose them
The adolescent movers did not vanish from the estimate. They were averaged into it. The abstract puts it in one sentence: “Moving as an adolescent has slightly negative impacts, perhaps because of disruption effects.”
Any mean effect pools people who differ in four ways: how much risk they carried to begin with, how strongly they respond to the treatment, what it costs them to take it up, and whether it reaches them at all. Once those are pooled, the mean reports all four at once and distinguishes none of them.
Distributional analysis is the habit of asking the estimate to come apart again. What happened at the bottom of the outcome distribution rather than at its middle. What happened in the tail, where the losses live. What the effect looked like inside groups fixed before anyone saw a result — age at move, in the case above. And, since the policy spent something real, where the resources moved.
That work is still estimation. It produces more numbers, and none of them says what to do. Deciding that a small decline for one group is acceptable because the total went up is a value judgment. Deciding it is unacceptable is also one. Whose outcomes count, which harms are ruled out whatever they buy, how gains should be spread across people who did not start level: somebody has to make those choices and defend them. Causal inference tells you what each choice would cost. It does not tell you which cost to accept.
This is why a single fairness number never closes the argument. At least three questions hide inside the word, and each has its own documented failure.
Are the groups equally offered the treatment, and equally likely to receive it once offered? A commercial risk algorithm affecting millions of US patients assigned Black and White patients the same scores at different levels of sickness. Obermeyer and colleagues reported it in Science in 2019. Their abstract gives the size of it: “Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%.” The failing quantity there was access to the programme, not the effect of the programme.
Do the causal gains differ across the groups that do receive it? In the RECOVERY trial, what dexamethasone did depended on how much respiratory support the patient already needed. Preliminary results were announced on 16 June 2020: 2,104 patients randomised to dexamethasone against 4,321 randomised to usual care. The Chief Investigators' statement reported: “Dexamethasone reduced deaths by one-third in ventilated patients (rate ratio 0.65 [95% confidence interval 0.48 to 0.88]; p=0.0003) and by one fifth in other patients receiving oxygen only (0.80 [0.67 to 0.96]; p=0.0021). There was no benefit among those patients who did not require respiratory support (1.22 [0.86 to 1.75]; p=0.14).” Treat those brackets as the announcement's own wording, not as the intervals to reuse. Six days later the same team's preliminary report gave 0.65, 0.51–0.82; 0.80, 0.70–0.92; and 1.22, 0.93–1.61. The final report in the New England Journal of Medicine gave 0.64 (95% CI 0.51–0.81), 0.82 (0.72–0.94) and 1.19 (0.91–1.55). Same drug, same trial, three pre-specified strata. In the third, the point estimate runs the wrong way.
And is any group carrying a loss you have already said you will not accept? That is the question the adolescent movers pose. No fairness figure quoted alone will show you which of the three you passed and which you failed.
An average over the Moving to Opportunity children can be precise to the dollar and still say nothing about the ones who moved after age 13.
Visual
The report runs decision first, estimate second
The order of a causal report is not decoration. Decision first, estimate second. A report built the other way round has no way to tell a finding from a rationalization written after the fact. ICH E9(R1) puts the estimand ahead of the analysis for exactly that reason. The target of estimation is fixed in advance, through its five attributes, so the question cannot quietly follow the answer.
So the map starts where the analyst should have started, by naming the decision that has to be made and the estimands that would inform it. Design is documented next: who was compared with whom, and under what identification claim. Then the estimates are reported, the average and the distribution it came from together — the pooled figure and the split by age at move on the same page. Then they are stressed, so a reader can see which conclusions survive when the assumptions are bent and which do not. Only at the end is the decision itself recorded, with its owners and the conditions that would end it.
Read the map at its joints rather than its boxes. Every arrow is a place where an assumption was accepted or a choice was made. Those are the places a reader should be able to push.
- 1
State decision and estimands
Average, subgroup, distributional, and policy targets.
- 2
Document design
Protocol, DAG, assignment, measurement, and support.
- 3
Report estimates
Point, interval, heterogeneity, and practical scale.
- 4
Show stress tests
Sensitivity, falsification, transport, and missingness.
- 5
Record decision
Benefits, harms, equity, owners, monitoring, and stop rules.
Key idea
A subgroup estimate published bare turns into a label
Breaking the average apart is the right instinct, and it is not free. A subgroup effect published without its uncertainty stops being an estimate almost immediately. It becomes a description of a kind of person — unresponsive, risky — and that description outlives the study that produced it. The damage is worst when the group was drawn along a social category rather than along anything in the mechanism. Worse again when the cell was so thin that the estimate was mostly noise.
Fine-grained tables carry a second hazard, and the agency that publishes the cells has measured it. In 2023 the US Census Bureau ran a simulated attack on its own 2010 Census. It reconstructed person-level records from just 34 of the 180 published table sets. The abstract states the result: “Using only published data, an attacker using our methods can verify that all records in 70% of all census blocks (97 million people) are perfectly reconstructed.” The study then correctly inferred race and ethnicity, with 95% accuracy, for 3.4 million people it classes as vulnerable population uniques — those whose race or ethnicity differs from the modal person on their block. The people most exposed by a published table are the ones rarest in it.
None of that argues for retreating to the single number, which is how the adolescent movers got lost in the first place. It argues for discipline about how the splitting is done. Fix the groups before seeing results, as RECOVERY did with its three respiratory-support strata, or fix them on a stated mechanism — a reason this treatment should work differently here. Publish uncertainty beside every subgroup figure. Refuse to report cells below a minimum size. Let the people being described review how they are described. And keep the language straight: an effect in a population is not a forecast about a person inside it. Writing it as though it were is exactly where the harm starts.
Transparency should expose evidence and uncertainty without turning subgroup estimates into identities.
Analogy
A levee blown at 10:03 pm, and 130,000 acres that took the water
At 10:03 pm on 2 May 2011 the US Army Corps of Engineers blew a hole in the Birds Point levee to relieve flooding at Cairo, Illinois. The river answered within the hour. The US Geological Survey records what followed: “The resulting inflow of water into the 130,000 acre floodway caused the Ohio river stage near Cairo to drop nearly 1/2 foot during the first hour of operation.” Roughly 130,000 acres of Missouri floodway — farmland and homes — took the water instead. The floodway had not been operated since 1937.
Average damage across the two places fell. It fell partly because the water had to go somewhere. The averaged figure is accurate. For the people inside the floodway it is also the description of a decision that was made about them, at a named hour, by a named agency.
This is the shape of the opening case: a positive mean produced partly by moving a burden rather than removing it. Water obeys a conservation law, so the transfer is easy to see and hard to deny. Here it was measured to the half-foot within the hour. An adolescent's lost earnings leave no waterline on anyone's wall. The arithmetic still cannot say where the barrier belongs. It can only say what each answer would cost, and to whom.
Averaging across a transfer makes the transfer disappear from the number, not from the 130,000 acres.
Steps
One memo has to satisfy two readers who never read alike
The release document is read by someone checking whether the estimate identifies anything, and by someone who has to sign for the consequences. The first wants the design, the claims, and the places they break. The second wants to know who gains, who loses, and what happens when it goes wrong. A memo written only for the first gets waved through by people who could not have caught the problem. A memo written only for the second cannot be proved wrong.
Build it from four pieces, in the order the map sets out. The estimand card leads, and it is not a free-form paragraph. Use the five attributes ICH E9(R1) names: treatment condition, population, variable, intercurrent-event strategy, population-level summary. Everything after the card is an answer to that question and to no other. The assumption register follows, each claim carrying an owner, so that disagreement has somewhere to go. Then the estimates, average and distribution together, never the average alone. A Moving to Opportunity memo that reported only the pooled gain would have been true and useless. Then the support report, the honest list of everyone the data could not speak for. Then the decision memo, where the recommendation finally appears next to the alternatives it beat.
Write the equity part as an argument rather than a metric. Say which of the three questions the policy is being judged against: access, as in the gap from 17.7% to 46.5%; differential gain, as in RECOVERY's 0.65, 0.80 and 1.22; or a loss you have ruled out, as in the adolescent movers. Then say who decided that it would be judged that way.
- 1
Summarize the target
Decision, interventions, population, and estimands.
- 2
Display distribution
Average, quantiles, subgroup effects, and harms.
- 3
List assumptions
Design evidence, diagnostics, and residual threats.
- 4
Evaluate policy
Value, capacity, access, burden, and alternatives.
- 5
Set lifecycle controls
Monitoring, appeals, update triggers, and retirement.
A causal report should make disagreement productive
The instinct on release is to close the gaps and let the result look settled. That is the wrong target. A report nobody can argue with has not resolved the disagreement. It has only moved it somewhere that leaves no record.
Useful disagreement needs the seams left showing. Keep the empirical estimate, the value judgment and the implementation assumption in separate sentences. Then a reader who accepts the estimate can still reject the recommendation, and point to the sentence where the two of you parted. Then show the movement. Which conclusions flip when the target population changes, when the harm constraint is tightened, when a sensitivity parameter is pushed as far as anyone would defend. A conclusion that survives all of that has earned something. A conclusion that flips under a value many people hold is the most useful line in the report.
The decision has to be a record rather than an announcement. For some deployers that is now law rather than taste. Article 27 of the EU's 2024 AI Act obliges public-sector and certain private deployers of high-risk AI systems to perform a fundamental rights impact assessment before deployment. It must name the categories of persons and groups likely to be affected, the specific risks of harm to them, and the human-oversight measures. It must also name, in the words of Article 27(1)(f), “the measures to be taken in the case of the materialisation of those risks, including the arrangements for internal governance and complaint mechanisms.” The result is notified to the market surveillance authority.
So: name the owners. Name the affected groups, the ones the estimate ran against among them, and the channel through which they can come back. Say what data is kept and for how long. Say what is being monitored, and state in advance which reading sends the policy back for redesign or takes it out of service.
Anyone who accepts your estimate and still refuses your recommendation should be able to point at the sentence where you parted.
Key takeaways
- An average can rise while a group inside it is harmed. Moving to Opportunity children who moved below age 13 earned $3,477 more per year — 31% more, against a control mean of $11,270. Those who moved after age 13 had, if anything, negative long-term impacts.
- Causal inference can tell you what each equity choice would cost. It cannot make the choice for you. Neither ICH E9(R1)'s five estimand attributes nor a fairness figure will make it either.
- Being offered the treatment, gaining from it, and being harmed by it are three separate questions. Correcting the risk algorithm reported in Science would raise the share of Black patients flagged for extra care from 17.7% to 46.5%. RECOVERY's rate ratios of 0.65, 0.80 and 1.22 are about who gains once treated.
- A subgroup effect is only publishable with its support, its uncertainty, and language that keeps it from hardening into a label. The US Census Bureau's own attack rebuilt every record in 70% of census blocks — 97 million people — from 34 of the 180 published table sets.
- An estimand card built on ICH E9(R1)'s five attributes, an assumption register and a support report are what let someone audit a causal claim without having to ask the analyst.
- A decision record is unfinished until it names monitoring, a route back for the people affected, and the reading that stops the policy. Article 27(1)(f) of the EU AI Act requires those arrangements in writing before deployment.