Causal inference
Time-to-Event Outcomes, Censoring, and Competing Risks
Handle survival outcomes, hazards, risks, censoring, competing events, and treatment strategies with time-aware estimands.
By the end you can
- Distinguish survival probability, cumulative incidence, hazard, and restricted mean survival
- Explain independent censoring assumptions and inverse censoring weights
- Define competing-risk estimands rather than treating competing events as ordinary censoring
- Choose effect measures aligned with time and policy decisions
Example
One hazard ratio of 1.24, sitting on top of six that ran from 1.81 to 0.70
The Women's Health Initiative estrogen-plus-progestin trial followed over 16,000 women for an average of 5.2 years. For coronary heart disease it reported a single hazard ratio: 1.24. That is the number that travelled.
Its own Table 2 carried six. Year by year the ratios ran 1.81, 1.34, 1.27, 1.25, 1.45 and 0.70 — for years 1, 2, 3, 4, 5 and 6-or-more. Nobody had cheated. Nothing had been computed wrongly. The single 1.24 is an average of a quantity that starts near 1.8 and ends below 1, so its size depends on how long the trial happened to run. Hernán works that out: the average would have been about 1.8 had the trial stopped at one year, 1.7 at two years, 1.2 at five.
Then ask the plainer question the ratio never answers. Behind the year-1 hazard ratio of about 1.8 sit absolute one-year risks of roughly 0.49% on treatment against 0.28% on placebo. A ratio close to two, and two absolute risks both under half a percent.
“Endowing a HR with a causal interpretation is risky for 2 key reasons: the HR may change over time, and the HR has a built-in selection bias,” Hernán wrote in 2010. Both reasons are legible in that one table. The ratio moves. And after the first year the comparison runs among the women still event-free in each arm — no longer the same kind of group, once treatment has changed who is left.
Four quantities were available to describe this outcome. The summary carried one.
- Survival is the probability that a unit is still event-free past some moment you name — the curve everyone pictures when they picture this kind of analysis.
- The hazard is not a probability over a period at all. It is the rate at which the event is arriving right now, counted only among those who have not had it yet. That is how one treatment produces 1.81 in year 1 and 0.70 from year 6 onward.
- Cumulative incidence is the probability that this particular event has happened by a given time, in a world where the other events that could arrive first are still allowed to arrive.
- RMST, restricted mean survival time, is the event-free time a unit can expect to accumulate up to a horizon you choose. An amount of time, not a ratio — and the quantity Royston and Parmar propose instead of one.
Nothing is defined until someone starts the clock
None of that comparison can be settled until somebody writes down what was being timed. A time-to-event study has to fix five things before it has a subject: when the clock starts for each person, what counts as the event, how far out you intend to look, what happens to someone who stops being observed, and which other events can arrive first and end the story.
Only then does a treatment effect have a shape. It can be stated as a difference in risk, as a gap between survival curves, as a difference in expected event-free time, or as a contrast between hazards. These are not interchangeable renderings of one result.
Take the three that get reported most often. A risk difference at a horizon is the change in the probability that the event has occurred by then — the form a patient or a planner can act on directly. An RMST difference is the change in event-free time accumulated up to a horizon τ, turning the whole curve into a single amount. A hazard ratio compares instantaneous rates among those still event-free.
That last clause is where the trouble lives. A hazard ratio is computed among survivors. Once treatment has changed who survives, the two sets of survivors are no longer the same kind of group. The ratio is also noncollapsible — its value in the whole population is not an average of its values within subgroups — and it depends on when you look. Read as though it were a risk ratio, it will mislead. Which is exactly what happened above.
And the moving part is not an exotic case. Somebody counted. Patient-level data was pulled back out of the Kaplan-Meier curves in 152 phase III oncology trials covering 129,401 patients, and the proportional-hazards assumption was tested figure by figure. “Among 304 Kaplan-Meier figures, 75 (24.7%) exhibited evidence of DPHs, including eight of 14 (57%) KM pairs from immunotherapy trials.” Deviation from proportional hazards in roughly one published figure in four, and in a majority of the immunotherapy comparisons. The 2019 report also names what the deviations travelled with: immunotherapy (OR 4.29, 95% CI 1.11–16.6), metastatic populations (OR 3.18, 95% CI 1.26–8.05) and non-overall-survival endpoints (OR 3.23, 95% CI 1.79–5.88). A single ratio averaging something that keeps moving is the ordinary case, not the pathological one.
A Cox model hands you a hazard ratio whether or not a hazard ratio was the question you walked in with.
Key idea
Censoring a death is a claim about a world in which nobody dies
The fastest way to get a wrong answer here is to tell the software that death is censoring.
Suppose recurrence is the outcome. A patient who dies gets coded as censored at that date, in the same column as a patient who moved away and stopped answering the phone. Those two people are in completely different situations. One might recur tomorrow, unobserved. The other never can.
Censoring means the event could still happen and you simply would not see it. Applied to death, that assumption describes a world where death has been removed and everyone stays available to recur. It is sometimes a question worth asking. It is a hypothetical one, and it needs stronger assumptions than the analysis that produced it usually carries.
How often is it asked by accident? Somebody sampled. A hundred recently published studies containing at least one Kaplan-Meier analysis were drawn at random from prominent medical journals. “Forty-six studies (46%) contained Kaplan-Meier analyses susceptible to competing risk bias,” van Walraven and McAlister reported in 2016. Only 16 of those susceptible studies (34.8%) gave both the outcome counts and the competing-event counts. Those 16 were the only ones where the damage could be measured at all. Six of them (37.5%) had Kaplan-Meier risk estimates biased upward by 10% or more relative to the true risk. Nearly half the sample was exposed to the error. In the checkable minority, better than a third of them had it.
So say out loud which of three questions is yours: total risk of the event in the world as it is, with the competing events left standing; the cause-specific process taken on its own terms; or risk in a hypothetical world where the competing event has been eliminated. All three are defensible. Only one of them is what a reader will assume you meant.
Whether a competing event is censored or counted decides which world your estimate describes, and it is not a setting to leave at its default.
Analogy
A tournament with exits for several reasons
A league table times how long each team lasts before its first defeat. Two teams never lose. They run out of money in midseason and leave. Record those departures as ordinary interruptions — as though each club would have carried on playing exactly as it had been — and the competition you end up describing is not the one that was played.
The comparison only stretches so far. A club's finances sit largely apart from its form. In the studies where this matters, the rate of the event and the reason for leaving both move with the treatment and with everything that has happened to the person up to that point.
The exits are not missing data lying around the outcome. They are part of what the outcome is.
Who left, and why they left, belongs inside the definition of the outcome rather than in a footnote about attrition.
Example
Four words that get traded for one another, and should not be
Reports use these four as if they were alternative summaries of the same curve. They answer different questions. Confusing them is precisely how a favorable ratio ends up being reported as a lower risk.
The separation is not a matter of taste; it was formalised. A competing-risk model has to put covariate effects directly on the cumulative incidence function rather than on the cause-specific hazard, and in 1999 Fine and Gray built one: a proportional hazards model for the subdistribution of a competing risk. Their reason for needing a second model at all is one sentence: “Unfortunately, the cause-specific hazard function does not have a direct interpretation in terms of survival probabilities for the particular failure type.” Cause-specific and subdistribution analyses are two estimands, not two computations of one. They have been two since 1999.
- The survival function answers one question at a time — has the event not yet happened by moment t — and you have to name the moment before it says anything.
- The hazard lives on a different scale entirely. It is a rate among units still at risk, recomputed at every instant, with everyone who has already had the event taken out of the denominator. It has no direct reading in terms of survival probabilities for the failure type you care about.
- A competing risk is any event that prevents the one you are counting, or alters the process that produces it. It is a feature of the outcome, not a nuisance to be censored away — which is why Fine and Gray had to build a model for the subdistribution rather than reuse the one for the cause-specific hazard.
- RMST folds the entire curve up to your chosen horizon into one number. That is why it can be reported as an amount of event-free time, instead of as a ratio that needs a conditioning clause to interpret.
Steps
Settle the estimand before the model, because the model cannot settle it
Order matters here, and it is the reverse of the common one. Most survival analyses begin with a fitted model and end by reporting whatever quantity that model happens to print. Done that way, the estimand was chosen by the software.
Write the protocol first. Fix the clock, the event, the horizon, the treatment of every kind of exit, and the measure you will report — on paper, before anything is fitted.
This is not only good practice. In clinical trials it is adopted guidance. ICH E9(R1), the Addendum on Estimands and Sensitivity Analysis in Clinical Trials, was adopted on 20 November 2019 and came into effect in the EU on 30 July 2020. It defines five estimand attributes: treatment, population, variable or endpoint, handling of intercurrent events, and the population-level summary. Those are the same five decisions listed above, wearing regulatory names. And it fixes when they are due: “The clinical questions of interest and associated estimands should be specified at the initial stages of planning any clinical trial.”
The addendum also names five strategies for intercurrent events: treatment policy, hypothetical, composite variable, while-on-treatment and principal stratum. The three worlds of the previous section have official labels here. Censoring a death is the hypothetical strategy — the world in which the competing event has been removed. Choosing it on purpose, in the protocol, is a different act from arriving at it through a default column in a dataset.
This is the only moment when those choices are still cheap. Once results exist, each one of them looks like a decision taken in order to obtain a particular answer, and you will have no way to show that it was not.
- 1
Fix time zero
Eligibility, assignment, and follow-up start.
- 2
Define events
Primary, recurrent, competing, and composite outcomes.
- 3
Choose horizon
Decision-relevant follow-up and administrative end.
- 4
Specify censoring
Reasons, measured history, and weighting assumptions.
- 5
Select measures
Risk, cumulative incidence, RMST, or hazard with interpretation.
Visual
Five decisions, each one narrowing the next
Set time zero. This is the moment from which everyone is counted. It has to mean the same thing for every unit in the study, or the curves are averaging people who are at different points in their own histories.
Define the event. Which occurrence stops the clock, and, just as important, which occurrences do not.
Set the horizon: the point out to which you are willing to claim anything. Past it, the data thin and the curve is being drawn by a handful of remaining units. The horizon is also what fixes the number. The same trial reports an average hazard ratio of about 1.8 at one year, 1.7 at two years and 1.2 at five.
Define censoring. That means stating, for each way of leaving, whether the event could still have happened afterwards.
Choose the measure last, because by then the first four have ruled most of them out. A named horizon makes a risk difference or an RMST available. A decision to leave competing events standing in the world makes cumulative incidence the honest curve. Only after all of that does a hazard ratio become one option among several, instead of the default.
- 1
Set time zero
Align eligibility and treatment strategy.
- 2
Define event
Type, adjudication, recurrence, and competing events.
- 3
Set horizon
Clinical or operational time window.
- 4
Define censoring
Loss, administrative end, switching, or competing event.
- 5
Choose measure
Risk, survival, RMST, hazard, or cumulative incidence.
Example
The same decision, in four fields that never talk to each other
Whether to censor a competing event, fold it into a combined endpoint, or model it directly is one decision wearing different vocabulary. In each of these, something ends the story early and the analyst has to say what that ending means.
One worked case shows the size of the gap. The EFFECT-HF cohort covered 16,237 Ontario patients hospitalised with heart failure. Within five years, 10,215 of them (63%) had died. Of those deaths, 5,970 were from cardiovascular causes — 58% of the deaths, 36.8% of the cohort. Yet the complement of the Kaplan-Meier estimate put the five-year incidence of cardiovascular death at 43.0%. “The use of the Kaplan-Meier survival function results in estimates of incidence that are biased upward, regardless of whether the competing events are independent of one another,” wrote Austin and colleagues in 2016, Fine among them. In the same data, cancer carried a subdistribution hazard ratio of 0.82 for cardiac death and a cause-specific hazard ratio of 0.96. One variable, one dataset, two numbers. Two different questions were being asked. Their rule of thumb: competing events above 10% in absolute terms merit serious consideration.
- In cancer follow-up, a patient who dies can never be observed to recur, so coding those deaths as censoring quietly asks what recurrence would look like if nobody died.
- A machine retired from the fleet cannot go on to fail in the way the study is counting — retirement removes the original failure mode rather than merely hiding it from view.
- When an account is closed, the later engagement events a churn model is waiting for have stopped being possible, not just stopped being observed.
- In a hospital, discharge competes with in-hospital deterioration. The two are opposite ends of the same stay, and treating one as censoring for the other describes patients who neither recover nor leave — the same structure that turned 36.8% into 43.0% in the heart-failure cohort.
Report the horizon and the world, not the ratio on its own
Lead with what a reader can act on: absolute risks, or the survival curves themselves, read at horizons that correspond to the decision in front of them. Beside 1.24, the pair 0.49% and 0.28% is the part a patient can weigh.
RMST belongs here too, and not only as a preference. Royston and Parmar set it out in 2013 as the area under the survival curve up to a horizon t*. They ended with a verdict rather than a suggestion: “We conclude that the hazard ratio cannot be recommended as a general measure of the treatment effect in a randomized controlled trial, nor is it always appropriate when designing a trial.” The design half of that matters, because trials are sized around the ratio too. A standard log-rank sample-size calculation needs about 510 events for 90% power against a target hazard ratio of 0.75 at two-sided 5%. The whole study gets built around the summary they decline to recommend.
Hazards still have a job. They describe how the rate of events moves across the follow-up — exactly what a single 1.24 concealed and the six period-specific ratios showed. Use them for that, and state plainly that they are computed among those still event-free.
Then show the exits. Break censoring patterns and competing events out by treatment group, so a reader can see whether the two groups were observed on the same terms. Report the competing-event counts themselves: in that sample of 100 published Kaplan-Meier analyses, only 16 of the susceptible studies gave a reader enough to check the bias at all.
And where leaving looks related to how the person was doing, a sensitivity analysis is not decoration. It is the only evidence anyone will ever have that the censoring assumption was checked rather than assumed.
A reader who cannot tell from your table which horizon and which world you meant will fill both in themselves.
Key takeaways
- A time-to-event study is only defined once the clock starts at the same moment for everyone and the event is written down. ICH E9(R1) has required the same five attributes since November 2019: treatment, population, variable or endpoint, handling of intercurrent events, and the population-level summary.
- Hazards, absolute risks and restricted means are three different causal summaries, not three renderings of one. Royston and Parmar conclude that the hazard ratio “cannot be recommended as a general measure of the treatment effect in a randomized controlled trial”.
- A hazard ratio is conditional on being event-free at that instant, which is why it is not a risk ratio. In Hernán's words, “the HR has a built-in selection bias”.
- Censoring is a claim about the risk faced by the people you stopped observing. That claim went unexamined in 46 of 100 randomly sampled published Kaplan-Meier analyses, and it overstated risk by 10% or more in 6 of the 16 where it could be checked.
- How a competing event is handled changes the quantity being estimated, not merely the precision of it. In one cohort, cancer carried a subdistribution hazard ratio of 0.82 for cardiac death and a cause-specific hazard ratio of 0.96.
- Absolute, time-specific measures usually match the decision better than a ratio reported without a horizon. The same trial's average hazard ratio would have read about 1.8 at one year, 1.7 at two and 1.2 at five.