Skip to content
AI.info

Causal inference

Propensity Scores and Covariate Balancing

Use propensity scores as balancing scores, choose covariates causally, and diagnose overlap and balance before outcome analysis.

By the end you can

Example

The model that predicted treatment almost perfectly

Propensity-score papers like to report how well their model predicts treatment. Two large reviews of the literature covered 224 articles between them. Ninety-one of those articles — 41% — reported a c-statistic for the propensity model. In each review, several reported one above 0.90. Five epidemiologists counted this up in 2011. The number was being printed the way a team prints its best diagnostic: as evidence that the adjustment had worked.

Their verdict is flat: “a high c-statistic in the propensity score model is neither necessary nor sufficient for the control of confounding”.

Set a thought experiment against those 0.90s. Take a randomised trial, the one design in which confounding bias is zero by construction. Build the propensity model out of its perfectly balanced outcome risk factors. The c-statistic comes back at 0.5. The model cannot tell a treated patient from an untreated one better than a coin toss, and there is no confounding bias left for it to control. Separation and confounding control are not the same axis. A score on one is not a report about the other.

Push from the other end and the mechanism shows. Chasing the c-statistic pulls in covariates strongly related to treatment but unrelated to the outcome. Scores pile up near zero for the untreated and near one for the treated. What remains, in the authors' words, is “relatively little overlap” between the two score distributions. Hardly anyone sits in the middle. Hardly anyone, that is, whose treatment could plausibly have gone either way. A study reporting a c-statistic above 0.90 has measured how far apart its two groups are. It has filed that measurement under success.

  • The propensity score is one number per person: the probability that this person would have received the treatment, given everything measured about them before treatment began.
  • It is called a balancing score because, when the assumptions hold, comparing people who share a score makes the measured covariates line up across the two groups — one number standing in for the whole list.
  • Overlap is the condition that makes any of it possible. Among people who look comparable, both treated and untreated cases have to actually exist in the data. That is precisely what “relatively little overlap” says has stopped being true.
  • Balance is something you check after the design has been applied, using standardized differences and side-by-side distributions. It is not something the fitted model can report about itself. A c-statistic of 0.90, or of 0.5, does not speak to it at all.

The score is a device for making groups comparable, not a classifier

Start from the assumption the whole method rests on. Conditional exchangeability says that among people who share the same measured baseline covariates, who got treated and who did not was, in effect, decided as if at random. Grant that, and the propensity score does something economical. It compresses a long list of baseline covariates into a single number. Comparing on that one number is enough to bring the covariates into line.

That economy is a theorem with a date. Rosenbaum and Rubin defined the propensity score in 1983, as the conditional probability of treatment assignment given the observed covariates. Then they summarised what they had proved in one sentence: “Both large and small sample theory show that adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates.” Read the last three words slowly. They are the guarantee's edge as well as its content. Sufficient — one scalar in place of the whole list. Observed covariates — and no others, ever. From there the score can drive matching, stratification, weighting, or adjustment inside a model.

None of that is treatment classification. A model tuned until its c-statistic stops improving has optimised the one quantity the theorem never asks for.

Which covariates enter the model is therefore a causal question, not a fit question. You include a variable because of where it sits in time and in the causal structure, not because it lifts the curve. And the result is judged after the design is applied, on two things. Did the groups end up comparable? Do comparable people exist on both sides?

Sorting travellers into lines with similar baseline characteristics before comparing two check-in policies gets the idea across. It is the line number doing the balancing work, not any individual traveller. The limit is worth holding onto. A line number cannot record why someone is in a hurry today. The score balances what was written down, and nothing else.

The propensity model has done its job when the two groups end up looking alike on paper — which is why a model that tells them apart flawlessly has done the reverse.

Example

Covariates that control confounding, and covariates that only sharpen the classifier

Go back to the covariate list, because that is where a c-statistic above 0.90 comes from. Keeping whatever predicts treatment strongly is exactly how a classifier is built. It is close to the opposite of how a confounding-control set is built. Two simulation studies in 2006 split the candidate variables into two kinds that no fit statistic can tell apart. Brookhart and colleagues state the consequence for anyone tuning a model: “standard model-building tools designed to create good predictive models of the exposure will not always lead to optimal PS models, particularly in small studies”. Each candidate variable has to be asked a different question. What does it cause, what causes it, and did it exist before treatment did?

  • A variable that strongly predicts who gets treated but reaches the outcome only through treatment will sharpen the classifier and control nothing. In the 2006 simulations, variables related to the exposure but not to the outcome increase the variance of the estimated exposure effect without decreasing bias. They also push scores toward the extremes. A 2011 study priced what conditioning on such an instrument costs. Across all its additive risk-difference scenarios, the largest absolute increase in bias from conditioning on the instrument Z was 0.018, on a crude bias of 0.141. Across the multiplicative scenarios the largest was 1.636, on a crude bias of 7.773. Small, next to the total estimation error. The advice follows from those magnitudes: “In these cases, minimizing unmeasured confounding should be the priority when selecting variables for adjustment, even at the risk of conditioning on IVs.”
  • A variable that predicts the outcome but barely predicts treatment looks useless to the classifier and should always be kept. In the same 2006 simulations, variables unrelated to the exposure but related to the outcome decrease the variance of the estimated exposure effect without increasing bias. It buys precision in the comparison you actually care about, and costs the c-statistic nothing you wanted.
  • Anything measured after treatment, on the path running from treatment to outcome, has no place in a baseline propensity model at all. Adjust for it and you subtract part of the effect you came to measure. No simulation of variance and bias applies here, because the quantity being estimated has changed.
  • A rough stand-in for a confounder nobody could measure directly is better than leaving it out. It still leaves confounding behind, roughly in proportion to how badly it stands in. That is the trade Myers and colleagues are pricing when they put unmeasured confounding ahead of the risk of conditioning on an instrument.

Comparison

One score, several designs, and each one answers about a different population

Having a score is not yet having a design. The same fitted numbers can be used four ways. Match people with similar scores. Sort everyone into strata and compare within them. Weight the sample so that the treated and untreated distributions are pulled toward each other. Or enter the score into an outcome model as a covariate.

The choice is not cosmetic, because each of these keeps a different set of people. Matching keeps the pairs it can find and sets the unmatchable aside. Weighting keeps everyone and lets the extreme scores speak loudest. The estimate that comes out at the end is an estimate for whoever survived the design.

Weighting is not one formula but a class. A 2018 paper made that explicit. It defined a whole class of balancing weights, and the familiar inverse-probability weights are one special case of it. Once weighting is a class rather than a formula, picking a member of it is picking a target population, and the analyst is doing the picking. So before touching an outcome, a team is choosing not only an estimator but the population their answer will describe.

FigureComparison · 3 columns

Matching

Construct comparable treated-control sets.

  • Transparent pairs
  • May discard units
  • Targets matched population

Weighting

Create a pseudo-population.

  • Marginal effects
  • Can amplify extremes
  • Target defined by weights

Stratification

Compare within score bands.

  • Simple diagnostics
  • Residual imbalance possible
  • Coarse approximation

Visual

The order of these steps is the part that protects you

The workflow runs in one direction and the direction matters. First choose the covariates, on causal grounds and on timing. Then fit the assignment model. Then apply the design — match, stratify, weight, or adjust. Then assess balance on the measured covariates. Then inspect support: look at where the scores actually landed, and whether comparable people exist on both sides.

The fourth step is the one that goes missing in practice, and there is a count of how often. An appraisal went through 47 medical articles published between 1996 and 2003 that used propensity-score matching. Austin's finding: “We found that only two of the articles reported the balance of baseline characteristics between treated and untreated subjects in the matched sample and used correct statistical methods to assess the degree of imbalance.” Two out of 47. Thirteen of the 47 — 28 per cent — used analysis methods appropriate to matched data.

Read the map for its assumptions and its decision points. The assumption sits at the top, in the covariate choice. The decisions sit at the bottom, where balance and support either pass or send you back. A reported c-statistic above 0.90 is the second step's output being read as though it were the fourth and the fifth, by an analysis that never performed either.

FigureProcess · 5 steps
  1. 1

    Choose covariates

    Use pre-treatment common causes and prognostic variables.

  2. 2

    Fit assignment model

    Estimate treatment probability without outcome-driven selection.

  3. 3

    Apply design

    Matching, stratification, weighting, or overlap targeting.

  4. 4

    Assess balance

    Compare means, distributions, interactions, and tails.

  5. 5

    Inspect support

    Scores, weights, and effective sample size by treatment.

Steps

Audit the design before you have seen a single outcome

Everything up to this point can be done with the outcome column hidden, and it should be. One of the two authors of the 1983 theorem made that his title in 2008: for objective causal inference, design trumps analysis. Rubin states the thesis in one line: “The thesis here is that observational studies have to be carefully designed to approximate randomized experiments, in particular, without examining any final outcome data.” Balance and support are criteria the design has to satisfy on its own terms, before any effect is estimated. Once you have seen the effect you can no longer honestly decide whether the design was acceptable.

So work through it blind. Check that every covariate you adjusted for was recorded before treatment. Apply the design. Then compare the two groups on each covariate and see how far apart they still are. The post-design toolkit was set out in 2009: standardized differences, variance ratios, higher-order moments, quantile-quantile plots and side-by-side boxplots. The same paper records the threshold most people compare against, along with its actual status: “While there is no clear consensus on this issue, some researchers have proposed that a standardized difference of 0.1 (10 per cent) denotes meaningful imbalance in the baseline covariate.” A convention, then, not a law. Worth quoting as one.

Then look at the distribution of scores in both arms and find the regions where one arm has nobody. Note how many units the design discarded and who they were.

If the answer is that the design fails, that is a finding you can act on, and Rubin licenses the strongest version of acting on it. A candidate data set will often have to be rejected as inadequate. Either key covariates were never measured, or the distributions of key covariates do not overlap between the treatment and control groups. Careful propensity score analyses are exactly what reveals that. Find it after the effect estimate is on the table and you will be tempted to argue with it.

FigureProcess · 5 steps
  1. 1

    Specify causal covariates

    Use timing and DAG roles.

  2. 2

    Fit candidates

    Logistic, flexible, or balancing-focused models.

  3. 3

    Apply design

    Match, weight, or stratify under a target estimand.

  4. 4

    Check balance

    Means, variances, quantiles, interactions, and subgroup tails.

  5. 5

    Check information

    Overlap, weight concentration, and effective sample size.

Key idea

When nobody could have gone either way, there is nothing to compare

Scores crowded near zero and one mean the same thing in plain terms. For most people in this data, a comparable person under the other treatment does not exist. That is what “relatively little overlap” describes, and it is what a c-statistic above 0.90 is reporting. The consequences are mechanical. Weighting turns unstable, because a handful of units with extreme scores end up carrying the whole estimate. Matching either throws away most of the sample or starts accepting pairs that are not really alike.

There are honest responses. Redefine the target population to the region where both treatments genuinely occur. Or improve the design. Or state that the treatment policy in this data never gave some people a real chance of the other option, so their effect cannot be learned from it.

Redefining the population has a tool. Overlap weights are the member of the balancing-weights class in which each unit's weight is proportional to the probability of that unit being assigned to the opposite group. Emphasis then lands on the people who could plausibly have gone either way, and drains away from those who could not. That is not only an intuition about where the evidence lives. It is a property: “The overlap weights are bounded, and minimize the asymptotic variance of the weighted average treatment effect among the class of balancing weights.” Bounded is the word doing the practical work. Unbounded weights are what let a few extreme units run the estimate.

Near-total separation is a result in its own right: it says the data cannot answer the question in the form it was asked.

Take the simplest design that passes, then say what is still wrong with it

Given all of that, the rule for choosing is unglamorous. Take the simplest score and the simplest design that achieve acceptable balance on the measured covariates and adequate support for the population you mean to describe. A more elaborate assignment model earns its place only by improving those two diagnostics. Not by fitting better. And not by separating the groups more sharply — which, as Westreich and colleagues showed, is neither the thing you need nor evidence that you have it.

Then report what remains. Name the covariates whose standardized differences are still above the 0.1 convention after the design. Say how many units the design set aside. On the 1996–2003 record, doing this at all puts you in a minority of two.

And keep one limit in view, because it is the one the whole method cannot escape. Rosenbaum and Rubin's theorem removes bias due to all observed covariates. Every diagnostic in this lesson looks at variables that were measured. A design that balances all of them perfectly says nothing whatsoever about the ones nobody recorded.

The last paragraph of an honest propensity analysis is the one listing what the design never managed to balance.

Key takeaways