Skip to content
AI.info

Evaluation

Post-Deployment Evaluation, Drift, and Delayed Labels

Build monitoring that separates input drift, score drift, decision drift, performance change, label delay, and feedback-loop effects.

By the end you can

A changed score distribution is not yet a model failure

A fraud system’s average risk score rises after a new merchant category launches. The shift may reflect genuine risk, a benign change in traffic mix, a feature bug, or a policy change. Monitoring detects changes. Evaluation decides whether those changes matter for performance and decisions.

How far apart the two can sit is documented. The Epic Sepsis Model was proprietary and already widely implemented in hospitals when outsiders validated it, across 38,455 hospitalisations of 27,697 patients at Michigan Medicine. Wong and colleagues published the result in JAMA Internal Medicine in 2021. Sepsis occurred in 7% of those hospitalisations. The model raised an alert on 6,971 of them, 18%. By the dashboard, the system was working hard. Measured against outcomes, its area under the ROC curve was 0.63 (95% CI 0.62-0.64). At the recommended score threshold of 6, its sensitivity was 33%. Under “Conclusions and Relevance” the authors wrote: “This external validation cohort study suggests that the ESM has poor discrimination and calibration in predicting the onset of sepsis. The widespread adoption of the ESM despite its poor performance raises fundamental concerns about sepsis management on a national level.”

The finding held at other sites. Ostermayer and colleagues ran the same model over 145,885 encounters in two county emergency departments. Their 2024 report gives a sensitivity of 14.7% and a positive predictive value of 7.6% within a six-hour window. Alert volume was never the performance number. No amount of it substitutes for an outcome-anchored evaluation.

Drift is a symptom that starts an investigation, not a diagnosis by itself.

Example

The model changes which labels become visible

A risk model sends high-score applications to manual review. The review queue, not the world, is what the dashboard ends up describing. The bail literature calls this structure selective labels.

  • Reviewed cases: investigators examine them, so detailed outcomes exist — and this is the only population the dashboard silently reports on.
  • Low-score cases: many proceed without review, so hidden fraud may remain undiscovered, and an absent label is not a negative one.
  • Policy update: a lower threshold pulls a different mix of cases into review, so observed prevalence among reviewed cases moves with no change in the model.
  • Naive dashboard: apparent precision changes partly because the labeled population changed; the metric is describing the review policy as much as the scores.
  • Remedy: sample unreviewed cases, and record the policy version alongside every row of the evaluation table so a metric can be attributed to a regime.

Visual

Five kinds of change

Each layer needs different evidence, and the five do not fail together.

Input drift is movement in feature distributions, missingness, volume, or source composition. The search behaviour underneath GFT is input drift, changing partly because its owner kept changing it. Score drift is model outputs moving even when the decision policy is unchanged. Decision drift is thresholds, rules, overrides, or capacity changing the actions taken. The Epic Sepsis Model’s recommended score threshold of 6 lives in this layer, not in the model. Concept drift is a change in the relationship between inputs and outcomes. It is where a treatment protocol can hide: asthma predicting lower pneumonia risk is a relationship manufactured by care, not by disease. Outcome drift is a change in prevalence, severity, cost, or user behaviour over time.

The layers matter because the same alert means different things at each. A score-drift alarm answers with a rollback. An input-drift alarm answers with a data contract. A concept-drift alarm answers with a re-examination of what the label means.

FigureHierarchy · 5 levels
  • Input drift

    Feature distributions, missingness, volume, and source composition change.

    • Score drift

      Model outputs move even if the decision policy is unchanged.

      • Decision drift

        Thresholds, rules, overrides, or capacity change actions.

        • Concept drift

          The relationship between inputs and outcomes changes.

          • Outcome drift

            Prevalence, severity, cost, or user behavior changes over time.

Comparison

Labels arrive on several clocks

An evaluation system should name each clock.

Immediate proxies are fast and abundant: a click, a reviewer action, a system acknowledgment, a short-term behaviour. They are useful for incident detection. They can also be gamed, or induced by the treatment itself.

Delayed outcomes sit closer to the real objective: chargeback, readmission, failure, retention, adjudicated quality. They arrive slowly. They can be censored. They are where performance truth lives. Sepsis in 7% of hospitalisations is a delayed-clock quantity. It is the clock on which an area under the ROC curve of 0.63 and a sensitivity of 33% were measured. The 18% alert rate runs on a different clock entirely and answers a different question.

The third clock never strikes. Never-observed outcomes belong to cases that were not reviewed, treated, or followed. A defendant the judge declined to release generates no failure-to-appear record at all. These create selective labels, cannot be assumed negative, and need sampling or causal design. They put a hard ceiling on what monitoring is entitled to claim.

FigureComparison · 3 columns

Immediate proxy

Click, reviewer action, system acknowledgment, or short-term behavior.

  • Fast and abundant
  • Can be gamed
  • May be treatment-induced
  • Useful for incident detection

Delayed outcome

Chargeback, readmission, failure, retention, or adjudicated quality.

  • Closer to real objective
  • Arrives slowly
  • Can be censored
  • Needed for performance truth

Never-observed outcome

Cases not reviewed, treated, or followed lack reference labels.

  • Creates selective labels
  • Cannot be assumed negative
  • Needs sampling or causal design
  • Limits monitoring claims

Window design changes alert behavior

Short windows detect incidents quickly but produce noisy estimates. Long windows are stable but slow. Seasonal cycles, traffic volume, label delay, and repeated users should guide window length and baselines.

Use event-time alignment and backfill metrics when delayed outcomes mature. Mark provisional values so incomplete labels are not compared with finalized history.

The GFT count is a lesson in baselines as much as in windows. Overly high prevalence in 100 out of 108 weeks, from 21 August 2011 to 1 September 2013, is not a week-level spike that a short window would flag. It is a sustained level shift, visible only against a long baseline and a comparison series. A monitoring design that watches for jumps will not see a system that is simply high, week after week, for two years.

A monitoring metric needs a maturity definition as well as a time window.

Key idea

Interventions can alter the target

Successful prevention removes its own evidence. A model that prevents failures may reduce the observed number of failures among high-risk cases. If the evaluation ignores the intervention, the model can look poorly correlated with outcomes precisely because action changed them.

The documented instance is pneumonia. A rule-based model trained on real pneumonia data learned “HasAsthma(x) => LowerRisk(x)”. Asthmatic pneumonia patients were routinely admitted straight to intensive care, and the treatment they received erased the risk the model was supposed to detect. Caruana and colleagues wrote it up in 2015: “The good news is that the aggressive care received by asthmatic pneumonia patients was so effective that it lowered their risk of dying from pneumonia compared to the general population. The bad news is that because the prognosis for these patients is better than average, models trained on the data incorrectly learn that asthma lowers risk, when in fact asthmatics have much higher risk (if not hospitalized).”

Accuracy did not catch it. On the same data a neural network reached AUC 0.86, and would therefore have deemed asthmatics low risk. Eran Tal reconstructed the case in 2023. A strong held-out score certified a model that had learned an admission protocol and mistaken it for a prognosis.

Store recommended action, actual action, timing, and outcome. Separate predictive monitoring from policy-effect monitoring.

Observed outcomes after intervention are not untouched labels for the original risk.

Case

Bail decisions, where only the released defendant has an outcome

Whether a defendant fails to return for a court appearance is recorded only when a judge released them. Lakkaraju and colleagues named that condition in 2017: selectively labelled data, in which “the observed outcomes are themselves a consequence of the existing choices of the human decision-makers”. Their example is bail. “For instance, in the context of judicial bail decisions, we observe the outcome of whether a defendant fails to return for their court appearance only if the human judge decides to release the defendant on bail.” The policy chose the sample. A detained defendant contributes no row to the outcome column, and no volume of monitoring will grow one.

Their remedy declines to invent the missing rows. They propose “an approach called contraction which allows us to compare the performance of predictive models and human decision-makers without resorting to counterfactual inference”. It exploits the differences between individual judges. Some release more readily than others, so the same kind of defendant appears on both sides of the label boundary.

What rides on getting this right is measured in the same team’s companion paper of 2017, Human Decisions and Machine Predictions. Its policy simulation reports crime falling by up to 24.8% with no change in jailing rates, or jail populations falling by 42.0% with no increase in crime. Those are the numbers an evaluation built on the observed labels alone cannot produce, and cannot check.

Analogy

A river gauge downstream from a movable dam

The gauge downstream reads the weather and the dam operator at once. Operators move the dam whenever upstream risk rises, so the level on the chart is never only what the sky did.

The measured version of that picture is Oakland. The neighbourhoods PredPol’s algorithm selected in Oakland drug-crime records already carried about 200 times more drug-related arrests than the rest of the city. Lum and Isaac reported the consequence in Significance in 2016: “Using PredPol in Oakland, black people would be targeted by predictive policing at roughly twice the rate of whites.” Estimates from the 2011 National Survey on Drug Use and Health put drug use at a roughly uniform rate across the city. The arrest record was the dam, not the weather.

The mechanism was then proved rather than illustrated. In 2018 Ensign and colleagues built an urn model that shows the loop running away: discovered incidents only occur in neighbourhoods the algorithm already sent police to. A risk model that picks which cases get looked at is in the same position. The cases carrying outcomes are the cases it chose. The ones it passed over have no reading at all. Post-action outcomes cannot be read without the history of the actions taken.

The system helps create the data used to judge it.

Steps

Connect monitoring to an operational response

A dashboard without ownership does not manage risk. One regulator has written the ownership down. Section 3308 of the Food and Drug Omnibus Reform Act of 2022 added section 515C to the FD&C Act, now 21 U.S.C. 360e-4. A manufacturer may change an approved device without a supplemental application if the change is consistent with an approved predetermined change control plan. The plan is filed before the change, not after the incident.

FDA guidance on those plans for AI-enabled device software, issued in December 2024 and reissued in August 2025, asks the Modification Protocol to “identify any triggers for re-training (e.g., when the quantity of new data reaches a certain size or when a drift in data is observed over time)”. Drift there is not a discussion topic. It is a named trigger with a pre-declared response. On the monitoring duty itself the guidance states: “As part of a manufacturer's responsibility to ensure that devices remain safe and effective, FDA anticipates that manufacturers will monitor their device's safety (e.g., adverse events) and effectiveness (e.g., performance) over time as modifications are implemented consistent with their authorized PCCP.”

The operating loop is the same five moves whether or not a regulator is reading. Detect: monitor data contracts, features, scores, decisions, latency, and the outcomes actually available. Triage: classify the alert as instrumentation, population, model, policy, or label-process risk. Confirm: use sampled review, delayed labels, slices, and comparison with a control or baseline. Act: roll back, adjust policy, retrain, restrict scope, or collect evidence. Learn: turn incidents into regression tests, and update thresholds, owners, and runbooks.

FigureProcess · 5 steps
  1. 1. Detect

    Monitor data contracts, features, scores, decisions, latency, and available outcomes.

  2. 2. Triage

    Classify the alert as instrumentation, population, model, policy, or label-process risk.

  3. 3. Confirm

    Use sampled review, delayed labels, slices, and comparison with control or baseline.

  4. 4. Act

    Rollback, adjust policy, retrain, restrict scope, or collect evidence.

  5. 5. Learn

    Turn incidents into regression tests and update thresholds, owners, and runbooks.

Key takeaways