Skip to content
AI.info

Causal inference

Matching and Subclassification

Use exact, coarsened, Mahalanobis, propensity, and hybrid matching plus subclassification with design-aware inference.

By the end you can

Example

5,735 critically ill patients, and the question of who the answer belongs to

Nobody in this study was randomised. Between 1989 and 1994, five US teaching hospitals followed 5,735 critically ill adult ICU patients in nine prespecified disease categories. The question was whether a right heart catheter helped them or harmed them. With no randomisation, the design itself had to decide which patients could stand in for which.

That decision was made on background information alone. A propensity score for receiving right heart catheterization was built by multivariable logistic regression. Patients who received RHC were then case-matched to patients who did not. The number came back against the procedure. JAMA published it in 1996: “By case-matching analysis, patients with RHC had an increased 30-day mortality (odds ratio, 1.24; 95% confidence interval, 1.03-1.49).”

That is what a matched comparison buys. What it costs is written in the patients who could not be paired. The standard review of matching methods puts the cost in one sentence: “When estimating the ATT it may be fine (and in fact beneficial) to discard controls outside the range of the treated individuals, but discarding treated individuals may change the group for which the results apply (Crump et al., 2009).”

Read that asymmetry slowly, because it is the whole lesson. Dropping unusable controls costs you nothing you were trying to estimate. Dropping treated units costs you the question. Nothing is hidden and nothing is falsified. The target moves while everyone is busy improving the comparison.

  • Before any of it, somebody has to write down what counts as similar. In the RHC study that rule was a propensity score for treatment, estimated by multivariable logistic regression from the recorded baseline information. Most of the design sits in that one line.
  • A caliper is the point at which similar stops being close enough. Choosing where to put it is a judgement about how much mismatch you can live with.
  • A separate decision is whether one comparison patient may stand in for several treated ones, or is used once and then spent.
  • Once the unmatched treated patients are set aside, the group the estimate speaks for is no longer the group that was eligible for the study. Stuart puts that sentence in print. Most matched analyses leave it out.

Matching is a design decision, made before any outcome is seen

The timing is where the authority comes from, and it is not a matter of taste. The work splits into two stages with a wall between them. Stuart's 2010 review states the rule: “Stage (1) uses only background information on the individuals in the study, designing the nonexperimental study as would be a randomized experiment, without access to the outcome values. Matching methods are a key tool for stage (1). Only after stage (1) is finished does stage (2) begin, comparing the outcomes of the treated and control individuals.”

The reason she gives for the wall is not tidiness. Working this way, she writes, “precludes the selection of a matched sample that leads to a desired result, or even the appearance of doing so.” Both halves of that clause carry weight. An analyst who can see the outcomes can shop for a caliper. An analyst who could have seen them cannot prove they did not.

The second payoff is that the answer leans less on a model. An outcome regression asked to compare a treated unit against controls who look nothing like it has to extrapolate. It states what would have happened in a region where it has seen almost nothing. Matching refuses that comparison instead of modelling it.

There is a family of ways to do the refusing: exact matching, coarsened exact matching, Mahalanobis distance, propensity matching, genetic matching. Each is a different opinion about what balance is worth and how many units you should lose to buy it.

What they share is a chain of commitments, best written out in the order you have to make them. Name the target: the effect among all treated units, among the ones you retained, or another population you can state. Set the hard constraints — the variables that must agree exactly, the pairs forbidden outright, the caliper past which nothing counts. Choose the distance. Decide how sets are built: how many controls per treated unit, whether controls may be reused, what happens when two candidates tie. Then audit balance, retention and pair distances. Only after that do you analyse outcomes, in a way that respects pairs, sets and reuse.

None of this touches what was never measured. A cause you never wrote down sits inside the estimate afterwards exactly as it sat there before.

What makes a matched study believable is that every comparison it refused can be pointed at.

Analogy

The renovated house and the one next door

Two houses stand on the same street, built the same year, roughly the same size. One was renovated and one was not. The gap between their prices is worth something, precisely because almost everything else about them is the same.

Then there is the mansion at the end of the road. It was renovated too, and there is nothing on the street to set beside it, so out it goes. The comparisons that remain are tighter. The estimate has also become a statement about ordinary houses rather than about every renovated house. The mansion is a treated unit, and dropping it is the move Stuart warns about. Not a loss of precision. A change of subject.

Pairing can only use what the register records: location, size, age. Whatever pushed an owner to renovate and also lifted the price, and was never written down anywhere, stays inside the answer however neatly the pairs line up.

Comparability is bought with sample, and the price is written in the list of houses that left the study.

Example

Each rule for building the sets moves the answer somewhere else

The trade from the opening case is not a single dial. Every construction rule shifts balance, precision and the population represented at the same time, and shifts them in different directions.

For one of those rules the literature has stopped waving at the trade-off and published a setting. That rule is the caliper. Take the linear propensity score, and suppose its variance in the treated group is twice what it is in the controls. Then, Stuart's review reports, “a caliper of 0.2 standard deviations removes 98% of the bias in a normally distributed covariate”. The standing advice is wider: “Rosenbaum and Rubin (1985b) generally suggest a caliper of 0.25 standard deviations of the linear propensity score.” Two numbers, one dial.

  • Pairing each treated unit with a single control is the version anyone can follow in a table, and it leaves most of the available comparison units unused.
  • Letting some treated units take several controls, where several good ones exist, buys precision unevenly. That is at least honest about where the data is thick and where it is thin.
  • Allowing a control to be reused gets everybody a closer partner, at the cost of a few comparison units carrying a large share of the estimate.
  • Tightening the caliper protects the pairs you keep and lengthens the list of treated units left without a partner. That is the opening trade-off with a dial on it, and there is a number on the dial. Austin's 2011 Monte Carlo study concluded: “When estimating differences in means or risk differences, we recommend that researchers match on the logit of the propensity score using calipers of width equal to 0.2 of the standard deviation of the logit of the propensity score.” Where at least some covariates were continuous, that value, or one close to it, minimised the mean square error of the estimated treatment effect. It also eliminated at least 98% of the bias in the crude estimator.

Comparison

Each method is a bet about what balance is worth

No method dominates for every covariate mix and estimand, because they are not optimising the same thing.

Exact matching is the most transparent rule available. The pairs agree on the variables, and anyone can check it. Past a handful of covariates, though, almost nothing agrees on all of them and retention collapses. Mahalanobis distance works on the covariates themselves and stays interpretable, then loses its ability to discriminate as the covariate list grows. Propensity matching sidesteps the dimension problem by pairing on one estimated number. That is exactly why the balance it produces has to be checked rather than assumed. Genetic matching searches for the weighting that leaves covariates best balanced and pays for it in transparency. The rule that produced the pairs is now the output of a search.

Subclassification makes the cheapest bet of the lot, and it rests on the oldest number in this lesson. Five subclasses on a continuous covariate remove almost all of the bias it carries. W. G. Cochran worked that out in 1968, on age in the smoking–lung cancer data, and Stuart's review reports it: “Cochran (1968) provides analytic expressions for the bias reduction possible using subclassification on a univariate continuous covariate; using just five subclasses removes at least 90% of the initial bias due to that covariate.” That is where the 5–10 subclass convention comes from. The result was later carried over to quintiles of the propensity score, and Austin states the modern form: “stratifying on the quintiles of the propensity score eliminates approximately 90% of the bias due to measured confounders when estimating a linear treatment effect”. Five bins, nine tenths of the bias, no pairs to defend.

Coarsened exact matching makes the most explicit bet of all, because it is stated as a guarantee rather than an average. In 2012 King and two co-authors sorted propensity score and Mahalanobis matching into the class they call “equal percent bias reducing” (EPBR). Those methods guarantee no level of imbalance reduction in any given data set. Their properties hold only on average across samples, and only under unverifiable assumptions about how the data were generated. Against that they set coarsened exact matching, a member of the Monotonic Imbalance Bounding (MIB) class: “CEM and other MIB methods invert the process and thus guarantee that the imbalance between the matched treated and control groups will not be larger than the ex ante user choice.” The price of the guarantee is the coarsening itself. Values are grouped into bands, and mismatch inside a band is accepted.

FigureComparison · 3 columns

Exact/coarsened exact

Requires agreement on selected categories.

  • Transparent
  • Can discard many units
  • Depends on coarsening

Propensity matching

Matches treatment probabilities.

  • Reduces dimension
  • Needs covariate checks
  • Can pair different profiles

Mahalanobis/hybrid

Matches multivariate distance, often within calipers.

  • Local similarity
  • Scaling-sensitive
  • Hard in high dimensions

Steps

Build the matches blind, then audit what you built

Sequence matters more than any single choice inside it, and Stuart's two stages are that sequence. Name the population you intend to speak about first, because every later decision either preserves it or quietly replaces it. Write the hard constraints next, while no outcome information can lean on them. That includes the caliper, where 0.2 or 0.25 standard deviations of the linear propensity score is a defensible starting point rather than a house rule. Pick the distance rule and build the sets. Only then turn and look at what the design produced: how balanced the covariates are, how many treated units survived, how far apart the pairs you kept actually are.

Outcomes come last, analysed in a way that respects pairs, sets and reuse. Stage (2) does not begin until stage (1) is finished. Go back and adjust the caliper after you have seen the effect estimate, and the value of the whole sequence is destroyed retroactively.

Then the report needs the part that is easiest to leave out: how many treated units found nobody, and who they were. The RHC authors reported an odds ratio of 1.24, with a 95% confidence interval of 1.03-1.49. They did not stop there. The same abstract carries the sensitivity analysis and the recommendation for what should happen next.

FigureProcess · 5 steps
  1. 1

    Declare the estimand

    State whose effect the matched sample should represent.

  2. 2

    Choose constraints

    Exact matches, calipers, and prohibited pairs.

  3. 3

    Compare designs

    Balance, retention, and pair quality across methods.

  4. 4

    Describe exclusions

    Profile unmatched treated and control units.

  5. 5

    Analyze by design

    Respect pairs, sets, replacement, and clustering.

Key idea

The software says matched; the two units may share nothing that matters

Two units can carry nearly the same propensity score and still differ on the covariates that decide the outcome. The score summarises how likely treatment was, not how the unit was doing, and two routes to the same likelihood can start from very different places.

This is not an occasional accident. It is the method working as specified. Gary King and Richard Nielsen made exactly that argument in 2019, and their paper opens on it: “We show that propensity score matching (PSM), an enormously popular method of preprocessing data for causal inference, often accomplishes the opposite of its intended goal—thus increasing imbalance, inefficiency, model dependence, and bias.” The mechanism is specific. Once the data are balanced enough to approximate complete randomisation, PSM approximates random matching. Random matching raises imbalance even relative to the original unmatched data. The tool keeps pruning after the point where pruning has stopped helping.

Averages conceal this. A balance table can look reassuring overall while the worst pairs sit out in the tails, or inside one subgroup the average drowns.

So read the distributions rather than the summary line: the pair distances, which controls were reused and how often, where the two groups overlap at all, and which treated units were left out. A success status from the algorithm reports that it finished. It does not report that it found anybody comparable.

A matched pair is a claim the analyst is making, not a finding the algorithm handed over.

When the narrower claim is the stronger one

Set against the analysis it replaces, a matched design offers less imbalance on the measured covariates and far less dependence on the shape of an outcome model — for the units it retained. That last clause is the whole deal.

When treated units central to the question have no counterpart, the honest report says their effects are unsupported. Forcing them into poor matches does not recover them. It only moves the damage out of the methods section and into the estimate.

So describe how the target changed and which units left. That is the sentence the opening case needed. Then set a sensitivity analysis beside the number. The question is how strong something nobody measured would have to be to overturn this, and it does not have to stay rhetorical. The 1996 RHC paper answers it in one sentence, with a figure: “Sensitivity analysis suggested that a missing covariate would have to increase the risk of death 6-fold and the risk of RHC 6-fold for a true beneficial effect of RHC to be misrepresented as harmful.” Two six-folds, on both arms of the confounding path, printed next to an odds ratio of 1.24. A reader can now argue about whether such a covariate is plausible in an ICU. That is a far better argument than arguing about whether confounding exists.

The same authors declined the two easy endings. They did not declare the procedure harmful and they did not simply demand a trial. The findings, they wrote, “should be confirmed in other observational studies”, and “These findings justify reconsideration of a randomized controlled trial of RHC and may guide patient selection for such a study.” Matching earned them a number worth taking seriously and a shortlist of who to randomise. It never claimed to earn more.

Answer the smaller question you can defend, and say out loud which question you stopped answering.

Key takeaways