Skip to content
AI.info

Causal inference

Overlap, Positivity, Trimming, and Target Redefinition

Diagnose structural and practical positivity, choose overlap populations, and report how trimming changes the estimand.

By the end you can

Example

The comparison the ward door had already forbidden

A history of asthma makes pneumonia more dangerous, not less. A model trained on hospital records said the opposite, confidently, and the fault was not in the fitting.

In the mid-1990s Cost-Effective HealthCare study, a rule-based system trained on one of the pneumonia datasets learned that a history of asthma lowers the risk of dying of pneumonia. It wrote the rule down: “HasAsthma(x) ⇒ LowerRisk(x)”. Caruana and colleagues wrote the episode up in 2015, and they are exact about where the rule came from: “But it reflected a true pattern in the training data: patients with a history of asthma who presented with pneumonia usually were admitted not only to the hospital but directly to the ICU (Intensive Care Unit).”

So the contrast the model appeared to be reporting had been settled at the ward door, before any record was written. An asthmatic with pneumonia treated aggressively, against an asthmatic with pneumonia not treated aggressively — one arm of that barely exists in the data. Admitting practice made it barely exist in the world.

The dataset was not thin. The pneumonia data covers 14,199 patients: 9,847 train, 4,352 test, 46 features, 1,542 deaths. That is a death rate of 10.86%. The train/test folds come from Cooper and colleagues in 1997. Two decades after the original system, a GA2M model fitted to that same data rediscovered the same asthma pattern. Size was never the issue. Every additional record collected under the same admission practice reproduces the same hole in the same place.

The ideas below separate the part of this problem that data can fix from the part it cannot.

  • A strategy can be structurally impossible for part of the target, the way a course of care without immediate ICU admission was effectively unavailable to an asthmatic arriving with pneumonia.
  • A strategy can be permitted and still almost never chosen, which looks similar in a table and is a different problem underneath.
  • The overlap population is what you have left once you keep only the units that could credibly have gone either way.
  • Redefining the target means moving the population or the contrast until the question matches something the data could ever have shown.

Positivity is a question about what could have happened, not about sample size

Positivity is the requirement that every option in the comparison had some real chance of occurring for every kind of unit the answer is meant to cover. Where that chance is zero, there is nothing to estimate.

It fails in two ways, and the difference decides what to do next. Theoretical or structural violations arise because, in the words of Petersen and colleagues in 2012, “it may be theoretically impossible for individuals with certain covariate values to receive a given exposure of interest”. Practical violations are finite-sample. The option was allowed and simply did not happen much. The estimate then leans on a few unusual units and inherits their peculiarities.

What separates the two is what more records would do. Their paper says it in one line: “The threat to causal inference posed by such structural or theoretical violations of positivity does not improve with increasing sample size.”

Their worked HIV example shows the practical kind in miniature. It uses 401 treatment change episodes with 35 baseline covariates. The mutation p82AFST is present in 25% of episodes, p82MLC in only 1%. Nothing forbids that 1% cell. There is simply almost nobody in it. Any answer that reaches into it carries the quirks of a handful of episodes into a population-level claim.

Software will not stop you in either case. Fit the model and predictions appear for every covariate pattern handed to it. That includes the patterns where one arm is empty, and the patterns where it holds four people. Those numbers are extrapolation wearing the clothes of measurement.

What keeps this honest is a short sequence of commitments made out loud. Name the population and the strategies being compared. Ask where each strategy was possible in principle, before looking at any data. Measure the support that exists, using estimated propensities, counts inside local covariate cells, and the density of covariates in each arm. Choose a response. Then say which population the final number actually describes.

Trimming and overlap weighting sit on that list of responses. Both change the population represented. That is not a side effect of the method. It is the method, and it is the part that has to be written down.

Ask where each option was even possible before asking how well it worked.

Example

Impossibility hides inside a healthy-looking sample

A balanced-looking treated share across the whole sample tells you nothing about any slice of it. Split the same data by severity, by region, by account type, by period, and one arm can empty out inside a slice while the overall balance stays reassuring.

There is a proved version of this, and it points at the habit analysts reach for first. Adding covariates to make unconfoundedness believable makes overlap harder to satisfy. D'Amour and colleagues open their 2017 paper on high-dimensional settings with the tension itself: “Researchers often argue that unconfoundedness is more plausible when more covariates are included in the analysis. Less discussed is the fact that covariate overlap is more difficult to satisfy in this setting.” They do not leave it as a warning. They derive bounds on average covariate-mean imbalance that tighten as the dimension grows. The long covariate list added to make one assumption believable is the same list that strains the other.

These are the shapes it usually takes.

  • At high severity one treatment cannot be given safely, so the comparison ends exactly where the patients are sickest.
  • A service exists only in selected regions, and everywhere else the choice was never a choice.
  • A hospital sends asthmatic patients presenting with pneumonia straight to intensive care, which leaves that group with no other arm to be compared against.
  • A new technology appears only in recent calendar periods, so a question about earlier periods is a question about a treatment that did not yet exist.

Analogy

Two roads, and the stretch where only one was ever built

Two roads run between the same pair of towns, and the obvious way to compare them is to time the drive on each. Along the stretch where both were built, that works. Along the stretch where only one was ever laid, the difference in travel time is not small or noisy. There is nothing to subtract, and the map shows that at a glance.

Support behaves the same way, minus the map. Feasibility here can turn on an admitting practice or on a patient's condition rather than on terrain, so the missing stretch never draws itself. It has to be found by measuring where each option ever occurred.

Where one option was never built, the difference between the options is not small; it is undefined.

Example

These words get blurred together, and each one points to a different repair

Keep these terms apart. Each points to a different next step, and blurring them is how a design problem gets mistaken for a data problem. The first two names are Petersen and colleagues'. The last two describe what an analyst does about them.

  • Structural positivity — their theoretical violation — fails when an option is impossible for part of the target, and the only honest moves are to change the target or to bound what cannot be seen.
  • Practical positivity fails when both options were possible but one was rarely taken, the way a covariate cell holding 1% of episodes is legal and nearly empty, which is a variance and extrapolation problem before it is anything else.
  • Trimming is the act of dropping units that fall outside a stated support rule, and it is defensible exactly as far as that rule is.
  • The overlap population is what survives the rule, the units with a substantial chance of either treatment, and it is a population you have to be willing to name out loud.

Steps

Put the excluded population on the first page

The report is where this either becomes a decision or quietly disappears. Write it in the order the analysis actually ran.

Naming the population is not house style. In clinical work it is a regulatory instrument. ICH E9(R1), the addendum on estimands, was adopted in November 2019 and took effect on 30 July 2020. Its glossary defines an estimand as “A precise description of the treatment effect reflecting the clinical question posed by the trial objective. It summarises at a population-level what the outcomes would be in the same patients under different treatment conditions being compared.” And it makes “The population of patients targeted by the clinical question” one of four attributes that have to be specified up front. Up front means before analysis, not in a limitations paragraph afterwards.

So: state the population and the contrast that were requested, in the requester's words, before touching either. Show where support exists and where it fails. Give the propensity distribution and the local counts that led you to say so. Say which units the support rule removed, and what those units have in common — a severity band, a region, a period. That is the sentence the decision owner needs. It is also the first one to be cut for length.

Then name the estimand you ended up with. If it differs from the one you were asked for, say so.

Hand this over before the effect estimates, not after. A reader who sees the number first will read everything around it as a defence of that number.

FigureProcess · 5 steps
  1. 1

    Plot support

    Propensity and covariate distributions by treatment.

  2. 2

    Identify structural gaps

    Policy, eligibility, and availability constraints.

  3. 3

    Quantify practical gaps

    Local counts, weights, and effective sample size.

  4. 4

    Compare target options

    Full population, treated, overlap, or restricted region.

  5. 5

    Document exclusions

    Who is lost and what decision remains unanswered.

Comparison

Every response to thin support sends the bill somewhere

No response to thin support is free. Each one buys something and charges for it somewhere else — in precision, in bias, or in the question now being answered.

Collecting more data is the only response that leaves the original question untouched, and it is useless against a structural failure. Under an unchanged practice, new records only repeat the impossibility. That is the point Petersen and colleagues make when they say the threat “does not improve with increasing sample size”. Restricting the sample to a supported region buys stability and changes the population. The estimate becomes trustworthy about fewer units.

Overlap weighting does something similar without a hard cut, and it has an exact definition and an exact payoff. Three statisticians proposed it in 2014 and published it in 2018: “We further propose a new weighting scheme, the overlap weights, in which each unit's weight is proportional to the probability of that unit being assigned to the opposite group. The overlap weights are bounded, and minimize the asymptotic variance of the weighted average treatment effect among the class of balancing weights.” That is the purchase: bounded weights, and minimum asymptotic variance inside that class. The charge is that the represented population is now defined by a weight function rather than by a sentence, and a weight function is hard to put in a report's first paragraph.

Bounds keep the original target and give up point identification. They return a range that is honest and often too wide to act on.

Redesign means a different contrast that is feasible for everyone. It costs the most before any analysis starts, and it is an established design method rather than an abstract suggestion. Hernán and Robins set it out in 2016: “Causal inference from large observational databases (big data) can be viewed as an attempt to emulate a randomized experiment—the target experiment or target trial—that would answer the question of interest.” Specifying that trial first — eligibility, strategies, assignment — is where an infeasible contrast gets caught. You have to write down who could have been randomised to what. It is the only response on this list that makes the original question answerable rather than merely answered.

FigureComparison · 3 columns

Collect more data

Increase observations under rare strategies.

  • Helps practical positivity
  • Cannot fix impossible policy
  • May be costly

Trim/restrict

Exclude unsupported units.

  • Improves comparability
  • Changes target
  • Needs transparent rule

Overlap weighting

Emphasize treatment-uncertain units.

  • Stable weights
  • Different estimand
  • Strong local support

Key idea

Turn the dial until the answer looks good and you have chosen the answer

The threshold is a dial, and the estimates move when it turns. That is the whole problem. An analyst who trims, looks, trims a little more and looks again is choosing a population by the answer it produces. The reported uncertainty has no way of knowing that any of it happened.

What blocks this is that the rule can be written from the propensity score alone. In 2009, in Biometrika, four authors characterised the optimal subsamples for which an average treatment effect can be estimated most precisely. Under some conditions the optimal selection rules depend solely on the propensity score. Their abstract gives the working version: “For a wide range of distributions, a good approximation to the optimal rule is provided by the simple rule of thumb to discard all units with estimated propensity scores outside the range [0.1,0.9].” The paper had circulated three years earlier under a blunter title: Moving the Goalposts: Addressing Limited Overlap in the Estimation of Average Treatment Effects by Changing the Estimand.

Two things there are worth separating. The rule touches only treatment assignment and baseline covariates, so it can be fixed while the outcomes are still out of sight. And the earlier title says out loud what applying it does: it moves the goalposts, by changing the estimand.

Set the support rule accordingly, and set it before any outcome is visible. Show what the estimate does across a range of plausible thresholds instead of at a favourite one. Report the excluded population before anyone reads an effect.

A threshold chosen after the estimates are on screen is a finding about the analyst, not about the population.

The smaller question is often the stronger result

When the full-population effect has no support, there are smaller questions with real answers. The effect among the treated. The effect on the overlap population. The effect inside a restricted and stated region of covariate space. Each corresponds to a contrast that could genuinely have gone either way, and each is a stronger scientific result than a full-population number propped up by extrapolation.

Trimming buys better numbers and a different quantity, and the people who studied it most carefully say both halves in one breath. Stürmer and colleagues simulated cohorts in 2010 with 8 covariates, 2 of them unmeasured strong risk factors. They trimmed the propensity score range asymmetrically. Trimming cut bias and mean squared error in most scenarios. Here is their own summary of what they had bought: “Treatment effect estimates based on PS range restrictions do not correspond to a causal parameter but may be less biased by such unmeasured confounding.”

That trade is worth making, and it carries one condition.

If the excluded group is the group the policy is about, narrowing is an evasion. There the work is to gather new evidence under a design that makes the alternative possible, or to state bounds and let them be as wide as they honestly are.

Whatever you estimate keeps its own name. Trimming away the asthmatic patients who arrived with pneumonia would not have answered the question about them. It would have stopped mentioning them, and the admitting practice that emptied one arm would have gone on doing so. A result on the overlap population answers a question about the overlap population. A report that lets a reader believe otherwise has done more harm than no report at all.

Renaming the estimand is honest work; leaving the old name on it is not.

Key takeaways