Skip to content
AI.info

MLOps

Model Quality Monitoring, Drift, and Delayed Labels

Monitor model behavior with leading indicators, delayed outcomes, slice analysis, drift diagnosis, and explicit uncertainty.

By the end you can

The model may be wrong for weeks before the label exists

A hospital readmission model produces scores today. The confirmed outcome arrives thirty days later. That thirty-day window is not a modelling convention someone chose. It is a payment rule.

Medicare judges hospitals on condition-specific 30-day risk-standardized unplanned readmission measures for six conditions: AMI, COPD, heart failure, pneumonia, CABG, and elective THA/TKA. The Hospital Readmissions Reduction Program behind those measures was created by section 1886(q) of the Social Security Act, and it began cutting payments on 1 October 2012, the first day of fiscal year 2013. The measures are calculated over a rolling performance period. CMS states the consequence on its own programme page: “The payment reduction is capped at 3 percent (that is, a payment adjustment factor of 0.97).”

That cap is not theoretical. In the fiscal year that began in October 2021, 2,499 hospitals were penalised, the average penalty was 0.64 percent, and 39 hospitals lost the full 3 percent. Those counts come from Jordan Rau's KFF analysis of the penalty data, “10 Years of Hospital Readmissions Penalties”.

So the team sees inputs, score distributions, overrides and workflow load immediately. True performance arrives thirty days later, for some patients only, and is then aggregated over a rolling window before it touches anything. Monitoring under delayed labels is an inference problem. Leading indicators tell the team where to look first. They do not certify degradation, and they do not certify safety.

Case

A deployed sepsis model missed two thirds of its cases

A widely implemented proprietary sepsis model was finally checked by someone other than its vendor. The cohort was large: “27 697 patients undergoing 38 455 hospitalizations”. Wong and colleagues published the external validation in JAMA Internal Medicine in June 2021, and the headline was a single number: “The ESM had a hospitalization-level AUC of 0.63 (95% CI, 0.62-0.64)”.

The number that matters to a ward is the one underneath it. “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” Two thirds of the sepsis cases missed. Nearly one stay in five alerted. A dashboard reporting only the area under the curve would have shown 0.63 and nothing about which two thirds the model lost.

Figure

Two thirds of the sepsis cases missed and 18% of stays alerted: the misses and the alert burden of a deployed model, side by side.

Visual

Evidence arrives on several timelines

A monitoring programme should state which claims are possible at each delay. The readmission measure sits in the third band, not the first. Thirty days pass before an outcome exists. A rolling performance period passes before it is aggregated. Only then does a payment adjustment apply, capped at 3 percent. Nothing observable on the day of the score can stand in for that.

FigureTimeline · 4 stops
  1. Immediate

    Service health, feature validity, model identity, scores, fallback, and workflow load.

  2. Short delay

    Human review, overrides, complaints, near-term proxies, and sampled adjudication.

  3. Outcome delay

    Confirmed labels, calibration, utility, subgroup performance, and missed cases.

  4. Long horizon

    Behavioral adaptation, feedback loops, long-term harm, and policy effects.

Example

A proxy improves while the real outcome worsens

A commercial risk-prediction algorithm ranked patients for enrolment in high-risk care management. Industry estimates put its reach at roughly 200 million people in the United States each year. An audit published in Science on 25 October 2019 found what it was really predicting: health care costs, used as a proxy for health need. Judged on the proxy, it worked. Obermeyer and colleagues state the result plainly in their abstract: “Thus, despite health care cost appearing to be an effective proxy for health by some measures of predictive accuracy, large racial biases arise.”

The size of the gap is the point. At the 97th-percentile threshold, correcting for that proxy would raise the fraction of Black patients automatically identified for the programme from 17.7% to 46.5%. On the same day the study appeared, the New York State Department of Financial Services and Department of Health wrote jointly to David S. Wichmann, CEO of UnitedHealth Group Incorporated, about Optum's Impact Pro. The letter told him that “neither you nor any other healthcare or insurance entity may produce, rely on, or promote an algorithm that has a discriminatory effect.”

  • Proxy signal: The algorithm predicted health care costs, and by some measures of predictive accuracy it predicted them effectively.
  • Hidden confounder: Cost and health need are not the same quantity, and the distance between them is not the same for every group of patients — which is where the large racial biases arise.
  • Delayed outcome: Measured against need rather than cost, the same 97th-percentile threshold should have been auto-identifying 46.5% Black patients, not 17.7%.
  • Incorrect response: Accuracy on the proxy was read as evidence the system worked, at a scale industry estimates put at roughly 200 million people in the United States each year.
  • Better design: Test the proxy against the outcome it stands in for, subgroup by subgroup, and expect the finding to come from outside — this one arrived as an external audit in Science and a regulator's letter on the same day.

Comparison

Monitoring signals support different claims

Treating every alert as “model drift” encourages the wrong fix. The three families of signal answer different questions, and only one of them is close to correctness. Even that one is not enough alone. A hospitalization-level AUC of 0.63 is an outcome-label statistic, and it still needed 1,709 missed patients and an 18% alert rate placed beside it before anyone could say what the model was doing to a ward.

FigureComparison · 3 columns

Data and score signals

Reveal distribution, coverage, freshness, and routing changes.

  • Fast and widely available
  • Useful for localization
  • Do not measure correctness directly
  • Can trigger investigation

Human and workflow signals

Reveal overrides, queue burden, complaints, and process adaptation.

  • Connect system to use
  • Can be biased by interface or incentives
  • Often available before labels
  • Need denominator and capacity context

Outcome labels

Measure task performance and utility when representative and mature.

  • Closest to correctness
  • May be delayed or censored
  • Can be selectively observed
  • Require cohort and label-quality analysis

Drift is evidence of change, not a diagnosis

Input drift may reflect seasonality, acquisition changes, source faults, or a legitimate population transition. Prediction shift can arise from input change, model routing, threshold change, or calibration. Outcome degradation can occur without large marginal feature shifts.

Google Flu Trends is the case where this was counted rather than asserted. It was built to predict CDC influenza-like-illness reports, and it drifted for two years while the model itself sat still. Lazer and colleagues wrote in Science in March 2014 that “GFT also missed by a very large margin in the 2011–2012 flu season and has missed high for 100 out of 108 weeks starting with August 2011”. By February 2013 it was predicting more than double the CDC's proportion of ILI doctor visits. They attribute the failure to “algorithm dynamics”. Google made 86 changes to search in June and July 2012 alone, so the inputs kept moving because the vendor kept changing the product that generated them. Their most deflating comparison is the baseline: even 3-week-old CDC data projected current flu prevalence better than GFT did.

The monitor should connect signals to a failure hypothesis and test the affected slices, time windows and decision pathways before choosing an intervention. A hundred misses in a hundred and eight weeks is a pattern. The diagnosis was upstream of the model.

Analogy

A farmer reads the field weeks before the harvest

A farmer walks the field weeks before harvest. Leaf color, soil moisture, pests and growth rate are useful leading indicators, but the final yield remains unknown. A greener field after extra fertilizer may still produce poor grain if another constraint dominates.

Inspecting a field changes nothing about which plants reach the harvest. A deployed model has the opposite property. It chooses which cases receive an action, and so chooses which outcomes ever become observable. A declined loan neither repays nor defaults. Its harvest is one nobody gets to weigh.

Predictive policing is where that property has been proved rather than intuited. A 2018 paper on runaway feedback loops opens with the mechanism: “Such systems have been shown susceptible to runaway feedback loops, where police are repeatedly sent back to the same neighborhoods regardless of the true crime rate.” A system retrained on the crime its own dispatches discover is reading its own harvest. Ensign and colleagues built a mathematical model of the process, proved why the loop occurs, and showed that resident-reported incidents attenuate it. They do not remove it. Removing it takes intervention. An earlier Oakland study by Lum and Isaac, of recorded drug crimes, found that targeted policing “would have been dispatched almost exclusively to lower income, minority neighborhoods”.

Leading indicators narrow hypotheses; mature outcomes decide whether the model delivered the intended result.

Key idea

Selective labels can make the monitor optimistic

When the system decides which cases receive an action, labels may exist only for the selected cases. Approval systems observe repayment for approved applicants. Alerting systems investigate only flagged events.

Naive performance estimates then describe the observed subset, not the intended population. Use exploration, external audits, counterfactual methods, or explicit scope limitations where appropriate.

The problem has a name and a paper behind it. Five researchers named the selective labels problem in 2017, and their example is judicial bail: “we observe the outcome of whether a defendant fails to return for their court appearance only if the human judge decides to release the defendant on bail”. The labelled cases are the ones the judge already chose, so the model is scored on a sample the incumbent decision-maker assembled. Their answer is a method rather than a caveat. They “develop an approach called contraction which allows us to compare the performance of predictive models and human decision-makers without resorting to counterfactual inference”.

A label is not representative merely because it eventually arrived.

Steps

Investigate a quality alert

Start with attribution and data validity before changing model parameters. The five steps below run from asset verification to the choice of intervention.

The last step is easier when its options were written down before the alert fired. FDA's final guidance on predetermined change control plans for AI-enabled device software functions was announced in December 2024 and reissued in August 2025. The notice of availability says what such a plan must contain: “This guidance recommends that a PCCP describe the planned AI-DSF modifications, the associated methodology to develop, validate, and implement those modifications, and an assessment of the impact of those modifications.” A manufacturer that has specified all three in advance can then make those changes without a new marketing submission for each one. Step 5 becomes a decision inside an agreed envelope, not an argument held under time pressure.

FigureProcess · 5 steps
  1. 1. Verify assets and inputs

    Confirm model, policy, features, population, and monitoring code.

  2. 2. Localize the change

    Analyze time, cohort, source, score band, fallback, and workflow slice.

  3. 3. Test hypotheses

    Distinguish fault, legitimate shift, policy effect, proxy change, and model degradation.

  4. 4. Wait or sample deliberately

    Use mature cohorts, adjudicated samples, or controlled exploration for labels.

  5. 5. Choose the intervention

    Repair data, change policy, recalibrate, retrain, contain, or continue observing.

Every alert should state the strongest justified claim

“Feature distribution changed” is a precise statement. “The model is bad” is not justified until the evidence connects the change to behavior or outcomes under the intended population.

Supervisors have written down the same discipline. Revised interagency guidance on model risk management was issued jointly by the Federal Reserve, the OCC and the FDIC on 17 April 2026, superseding earlier versions from 2011 and 2021. Its section on ongoing model monitoring defines the activity against a named list of causes: “Ongoing model monitoring involves an evaluation of the extent to which a model is performing as expected given potential changes in products, exposures, activities, clients, data relevance, or market conditions.” Where a model no longer performs as expected, the guidance says the situation “may warrant overlays, adjustment, or redevelopment”.

Read the scope before borrowing the authority. The guidance says it is expected to be most relevant to banking organisations with over $30 billion in total assets. And it is supervisory guidance rather than a binding rule: the agencies state that it does not set forth enforceable standards or prescriptive requirements. A monitoring runbook should say what each signal can and cannot establish, and what evidence is needed next. That includes being exact about what a cited authority actually requires.

Key takeaways