Skip to content
AI.info

Causal inference

Transportability and External Validity

Transport trial or observational effects across populations using effect modifiers, sampling weights, standardization, and structural assumptions.

By the end you can

Example

The same contract, two implementers, and an effect that vanished

Kenya ran its contract-teacher reform nationwide, and a randomized trial was built into it. The trial did something most trials cannot. It ran the identical contract-teacher intervention in parallel in all eight provinces and changed one thing between the arms: who implemented it. Where an international NGO ran the program, test scores rose by roughly 0.18 standard deviations. Where the Ministry of Education ran it, the effect was essentially zero.

The 2018 write-up says so in the first sentence of its abstract: “New teachers offered a fixed-term contract by an international NGO significantly raised student test scores, while teachers offered identical contracts by the Kenyan government produced zero impact.”

The contracts were identical. Same country, same period, same design. The authors trace the gap to implementation delays and to a differential interpretation of those identical contract terms. Not to teacher characteristics.

That matters for anyone tempted to fix the problem with weights. Hand an analyst the NGO arm and ask whether it carries to the Ministry. They could reweight the teachers until the demographic profiles matched exactly. They would learn nothing. Reweighting corrects the only difference it is capable of correcting: who the people were. What changed here was what the words fixed-term contract referred to once a different organisation was administering them.

Four terms describe that failure, and everything after this section depends on keeping them apart.

  • The source population is the group the evidence came from. Here that is the pupils and teachers in the arm an international NGO was running — the arm that produced the roughly 0.18 standard deviation estimate.
  • The target population is the group the decision is about, and it is usually the group nobody measured. Here, the schools a ministry would be running at national scale.
  • An effect modifier is anything whose value changes how much good the treatment does. In Kenya the modifier was the implementer. It moved the effect from roughly 0.18 standard deviations to zero without altering a word of the contract.
  • A transport formula takes the effect the source shows inside each slice of a modifier and reweights those slices to look like the target. It delivers a real answer only if the assumptions underneath it hold. Which is why no weighting scheme applied to teacher characteristics could have recovered the Ministry arm's result.

Example

Which differences between two populations actually change the answer

Not every imbalance between two populations does causal work. Source and target can differ in a dozen measured ways while the treatment goes on doing the same thing to both. Matching on everything visible is neither necessary nor sufficient. What matters is whether a difference sits on the mechanism the treatment runs through. And some of the differences that matter most were never recorded as variables at all.

One of them was measured on chest radiographs. Pneumonia prevalence was 34.2% at Mount Sinai, against 1.2% at NIH and 1.0% at Indiana University. Zech and colleagues trained pneumonia-detection CNNs across all three at once — 158,323 images, 112,120 from NIH, 42,396 from Mount Sinai, 3,807 from Indiana. Their 2018 paper reports what the baseline rates alone were worth: “The prevalence of pneumonia was high enough at MSH (34.2%) relative to NIH and IU (1.2% and 1.0%) that merely sorting by hospital system achieved an AUC of 0.861 (95% CI 0.855–0.866) on the joint MSH–NIH dataset.”

Nothing about the disease produced that 0.861. The site did. And the networks could read the site: they identified the source hospital system for 99.95% of NIH images and 99.98% of Mount Sinai images. The best internal model reached AUC 0.931. At the external site it scored 0.815.

Two of the four differences below separated the NGO arm from the Ministry arm in Kenya. Neither of them was a characteristic of a teacher.

  • How much good a treatment does in absolute terms depends on how badly the untreated would have fared. A population with less to lose has less to gain. A site where pneumonia prevalence is 34.2% is a different arithmetic problem from sites where it is 1.2% and 1.0%, before any model is fitted.
  • Most interventions are not self-contained. They lean on staff, administration and follow-up that the source had and the target may not. That is what implementation delays cost the Ministry arm in Kenya.
  • The same label can cover very different delivery. Identical contract terms, interpreted differently by two implementers, carried the same name into the report. One end produced roughly 0.18 standard deviations. The other produced zero.
  • Calendar time changes the comparison on its own. Standard care improves, pathogens and markets move, and other policies land on the same people while nobody adjusts for any of it.

External validity is a causal comparison between two populations

Strip the vocabulary away and the question is a comparison. Does what the treatment did to the people in the source also describe what it would do to the people in the target? That is not the same question as whether the source study was any good. A trial can be run impeccably — clean assignment, honest analysis, no shortcuts — and still be a statement about its own participants and nobody else. The two properties are separate. A strong result on the first says nothing about the second.

Differences between the populations start to matter when they change the size of the effect, who ends up treated, how the outcome is measured, whether people stay with the treatment, or which version of the treatment they receive.

Standardization and sampling weights address exactly one of those. It is worth watching the operation done properly, so the scope is unmistakable. ACTG 320 enrolled 1,156 US adults in 1996: 577 assigned to highly active antiretroviral therapy, 579 to a largely ineffective combination, followed for 52 weeks. Cole and Stuart standardized that result in 2010 to a different population — US people living with HIV in 2006, as estimated by the CDC. The effect survived the reweighting, attenuated by 12%. Their abstract states the condition attached to that number as plainly as the number itself: “Results from the trial apply, albeit muted by 12%, to the target population, under the assumption that the authors have measured and correctly modeled the determinants of selection that reflect heterogeneity in the treatment effect.”

That clause is the whole boundary. Weights line up the distributions of the variables somebody thought to record. They cannot conjure a modifier nobody recorded. They cannot turn two different interventions into one. They cannot repair an outcome that means something different in the target. And they cannot speak for a corner of the target that has no counterpart in the source at all.

Reweighting is the answer to a question that has to be settled before it: which differences here actually change the effect?

Visual

The audit that would have caught it, taken in order

The order matters, because each step decides whether the next one is worth doing.

Start by writing down the source and the target as two named groups of people rather than as a dataset and a deployment. Then compare the interventions themselves — not the labels, the actual service delivered at each end. Only after that go looking for the modifiers, the differences that plausibly change the size of the effect. Then ask the awkward follow-up: were they measured in the source at all?

Then assess support. For every kind of person the decision will cover, ask whether the source contains anyone like them. Transport last, and validate afterwards in the target.

This order is not a counsel of perfection. Two regulators already require the front half of it. ICH's harmonised guideline E17 covers multi-regional clinical trials. It was adopted in 2017, taken up by the EMA's CHMP the same year, and has been in effect since 2018. Its second basic principle is a deadline rather than an aspiration: “The intrinsic and extrinsic factors important to the drug development programme should be identified early.” Their potential impact is to be examined in the exploratory phases — before the confirmatory multi-regional trials are designed, not after they read out.

The third principle does the other half of the work. It writes the assumption down as an assumption. E17 states that “MRCTs are planned under the assumption that the treatment effect applies to the entire target population, particularly to the regions included in the trial”, and that “strategic allocation of the sample size to regions allows an evaluation of the extent to which this assumption holds”. The sentence a rollout usually leaves unspoken is printed here in the guideline, next to the design choice that lets somebody test it.

A team working in this order learns at the second step that the delivered service differs by implementer. That costs one meeting. Working in reverse, the same fact arrives after the money is spent. Which is precisely why the Kenyan team put the second implementer inside the trial instead of discovering it at national scale.

FigureProcess · 5 steps
  1. 1

    Define source and target

    Eligibility, setting, calendar time, and decision.

  2. 2

    Compare interventions

    Versions, delivery, adherence, and complementary resources.

  3. 3

    Identify modifiers

    Variables affecting treatment response and selection.

  4. 4

    Assess support

    Does the source cover target modifier combinations?

  5. 5

    Transport and validate

    Standardize, weight, triangulate, and monitor target outcomes.

Comparison

Generalizing and transporting are requests for different things

The two words are not used consistently. Some people reserve one of them for the case where the target is the wider population the study sample was drawn from, and the other for a target that is a genuinely separate population the study never touched. Others use them the other way round. Others use one term for both situations and let context decide.

Arguing about which label applies is wasted effort. Writing down a convention and then holding to it is not. Degtiar and Rose do exactly that at the top of their 2021 review: “Generalizability focuses on the setting where the study population is a subset of the target population of interest, while transportability addresses the setting where the study population is (at least partly) external to the target population.” Subset, or at least partly external. One sentence, and the ambiguity is gone for anyone reading that paper.

The same review supplies the other reason the argument is a distraction. It lists four distinct sources of external-validity bias — subject characteristics, setting, treatment, and outcomes — and notes that most generalizability and transportability methods address only the first of the four. The vocabulary dispute is over the name of the room. Three of the four ways the estimate can be wrong are outside the door.

What is never wasted is writing down, in a sentence anyone can check, which population produced the evidence and which population the decision is about. Most disputes over whether a result generalizes dissolve on contact with that sentence. The two people arguing turn out to have had different targets in mind the whole time.

FigureComparison · 3 columns

Trial generalization

Extend from trial participants to the trial-eligible population.

  • Selection into trial
  • Same setting
  • Sampling adjustment

Transportability

Move evidence to a distinct target population or setting.

  • More structural differences
  • Requires modifier theory
  • May change treatment version

Replication

Run a new study in target context.

  • Strongest direct evidence
  • Costly
  • Can reveal implementation change

Steps

Write the dossier before the rollout, not after it

A transportability dossier is a short document that compares the source and the target on four fronts and commits to an answer on each of them.

It starts with the mechanism — what the treatment actually works through — and asks whether that machinery is present in the target at all. It asks next whether the thing being delivered is the same thing, at the same intensity, by people with comparable training. That is the question the Kenyan trial answered by running two implementers. It asks whether the outcome means the same in both places and is recorded the same way. And it asks, for each kind of person the decision will cover, whether the source contains anyone resembling them.

None of this requires new data. It requires somebody willing to write no in a box.

The cost of the unwritten box has been measured. Opower's Home Energy Reports had been tested in 111 randomized trials, covering 8.6 million households at 58 US electric utilities as of February 2013. Hunt Allcott took the first ten sites and used them to predict the other 101. The prediction ran long. It overstated the true average treatment effect by 0.66 percentage points under a linear prediction, and by 0.41 percentage points under a weighted prediction, at p < 0.0001. The early sites were not a random draw of the eventual ones. Each of the first 11 had a frequency-adjusted effect of at least 1.34 percent, while 67 of the next 100 fell below that. His 2015 paper on site selection bias converts the gap into money: “Thus, in the context of a nationally-scaled program, these mispredictions would cause first-year retail electricity cost savings to be overstated by $560-920 million.”

Read the design of that failure carefully, because it flatters none of the usual defences. These were randomized trials, dozens of them, reweighted on rich microdata. The extrapolation still ran long by 0.41 to 0.66 percentage points. Internal validity was never the problem. Where the program had been run was.

FigureProcess · 5 steps
  1. 1

    Define populations

    Eligibility, setting, period, and sampling.

  2. 2

    Compare treatment versions

    Delivery, adherence, and co-interventions.

  3. 3

    List effect modifiers

    Measured, unmeasured, and plausibly structural.

  4. 4

    Assess overlap

    Covariate and implementation support.

  5. 5

    Plan target validation

    Pilot, outcome monitoring, replication, and stop rules.

Example

Four words that get used interchangeably and should not be

These arrive in the same meeting and get treated as near-synonyms. Holding them apart is most of the discipline. The last of them is the one that looks like it does the most work while doing the least.

The vocabulary is not house terminology. The transport formula got its formal statement in 2014, when Pearl and Bareinboim introduced selection diagrams and reduced transportability to do-calculus derivations. Those derivations decide, before any data are collected, whether a causal effect measured in the study population is recoverable in the target at all. Their abstract frames the whole enterprise as a permission rather than a procedure: “This paper treats a particular problem of generalizability, called "transportability," defined as a license to transfer causal effects learned in experimental studies to a new population, in which only observational studies can be conducted.”

A licence is granted or refused. That is the register the four words below belong in.

  • The source population is wherever the causal evidence was actually produced. It is a fact about the past and not adjustable — the NGO-run classrooms, the ACTG 320 enrollees of 1996, the first ten Opower sites.
  • The target population is wherever someone now has to make the causal decision. It is chosen by the decision, not by the data that happens to be available — US people living with HIV in 2006, the 101 later Opower sites, the schools a ministry will actually run.
  • A transport variable is a population difference you have to know about in order to connect the source effect to the target effect. It is exactly what selection diagrams mark. An unmeasured one quietly withholds the licence.
  • A sampling weight stretches or shrinks each source participant so that the participants collectively resemble the target. It aligns distributions and can yield a defensible number, as Cole and Stuart's 12% attenuation did. But only under their stated assumption that the determinants of selection were measured and correctly modeled.

Analogy

A crop result carried to a different climate

A yield measured on one farm does not carry to a farm in another climate because both planted the same seed. It carries, if it carries at all, once somebody has accounted for soil, rainfall and farming practice — the things the seed actually acts through. Match the two farms on the age of the farmer instead and you have matched something the harvest never depended on, then congratulated yourself on the balance table.

The seed is the fixed-term contract, identical in both arms in Kenya. The soil is who administered it, and how quickly.

The comparison has one limit worth carrying forward. A crop only responds to its environment. A program handed to an organisation can change the institutions and the behaviour around it, so the target does not sit still while it is being treated.

The farm and the classroom fail the same way — matched on the wrong variables, both look comparable right up until the harvest.

When a transported effect is still only a hypothesis

Somewhere in the target there is a kind of person the source study never enrolled. The estimate for them does not fail loudly. A weight grows enormous, or a fitted model extends its surface into a region where it has seen nothing, and a number comes out looking exactly like all the other numbers on the page.

No amount of standardization creates effect information for a combination nobody ever observed. And a tool can compute smoothly, in production, at scale, for years, while nobody has yet checked it against a target. The Epic Sepsis Model was proprietary and already deployed at hundreds of US hospitals when Michigan Medicine put it through an external validation. The study covered 27,697 patients and 38,455 hospitalizations, between 6 December 2018 and 20 October 2019. The model achieved an area under the ROC curve of 0.63 (95% CI 0.62-0.64). Of 2,552 septic patients it identified only 183 who were not already being treated promptly, or 7%. The 2021 validation reports the rest: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.”

Nothing errored. The alerts fired, in volume.

The honest responses are few and unglamorous. Look at the overlap between source and target before trusting anything that comes out. Narrow the target to the part the source can genuinely speak for. Go and collect bridging data on the part it cannot. Or report bounds and scenarios rather than one figure, and say plainly which of those scenarios the decision would still survive.

Transport the effect when the relevant modifiers were measured, when the source covers the target, when the treatments are comparable versions of the same thing, and when somebody has committed to validating the result in the target afterwards. Miss any of the four and what you are holding is a hypothesis. A reasonable prior for the next study, not a finding.

Which is why a small pilot in the target buys more than it appears to. Weighting can only surface the failures somebody measured, the ones Cole and Stuart's determinants of selection can reach. A pilot surfaces the ones nobody thought to record: the program delivered differently, as in the Ministry arm; the outcome recorded differently; the baseline rate that was 34.2% in one place and 1.2% in another; the specialist who was never there in the first place.

A number that computes without error on the target's data has demonstrated nothing whatsoever about the target.

Key takeaways