Skip to content
AI.info

Evaluation

Capstone: Design an Evaluation Program That Can Survive Contact With Reality

Design and defend an end-to-end evaluation program for a high-stakes AI-assisted service with incomplete labels, capacity limits, and changing conditions.

By the end you can

The system: urgent maintenance triage for municipal infrastructure

A city receives reports, sensor streams, photographs, and technician notes about water leaks, road damage, and electrical faults. An AI-assisted service ranks cases, predicts severity, drafts summaries, and recommends dispatch actions.

The city wants faster response without overlooking dangerous incidents, overloading crews, or disadvantaging neighborhoods with poorer sensors and fewer reports.

The brief is an exercise. The standard it is graded against is not. Every section that follows is anchored to a system that was actually built, actually deployed, and actually measured afterwards. A hospital sepsis alert. A health-plan risk score. A mammography reading aid. A driverless fleet. An air-traffic replacement. The capstone asks you to write, in advance, the evidence those programs produced late or not at all.

The capstone evaluates a socio-technical decision system, not a standalone model.

Visual

Break the proposal into testable claims

Each claim needs its own evidence. A system can pass one and fail another with no contradiction between the two results.

The Epic Sepsis Model carries two numbers, and both are true. At Michigan Medicine its hospitalization-level area under the curve was 0.63 (95% CI 0.62-0.64). Separately, its alerts identified 183 additional sepsis patients — 7% — beyond what timely clinical practice already caught. One is a prediction claim. The other is a service-outcome claim. Two questions about one deployed triage system, and neither answer substitutes for the other. Split the municipal proposal the same way, before any measurement begins.

FigureHierarchy · 5 levels
  • Triage ranking

    High-severity incidents appear early enough for dispatch capacity.

    • Severity prediction

      Risk estimates support thresholds, escalation, and abstention.

      • Summary quality

        Drafts preserve evidence, numbers, location, and uncertainty.

        • Tool safety

          Recommended dispatches use correct identifiers, permissions, and constraints.

          • Service outcome

            Response improves without unacceptable burden, delay, or geographic disparity.

Steps

Define units, populations, and label maturity — the label is the design decision

The program starts with the evidence contract. Most of that contract is a single choice: what will stand in for the thing you actually care about.

A widely deployed commercial health-risk algorithm, used to select patients for extra help, was not predicting illness. It was predicting health-care cost. Obermeyer and colleagues took it apart in Science in 2019. Cost has every property step 3 below rewards. It is recorded for everyone. It arrives without adjudication. It matures quickly. It is also not the outcome. At a given risk score, Black patients were considerably sicker. Correcting the target would have raised the share of Black patients referred for extra help from 17.7% to 46.5%. The paper closes on label construction rather than modelling: “We suggest that the choice of convenient, seemingly effective proxies for ground truth can be an important source of algorithmic bias in many contexts.”

Regulators moved the day the paper appeared. New York State's Department of Financial Services and Department of Health wrote jointly to UnitedHealth Group, opening an inquiry into the Optum Impact Pro algorithm the study described: “These discriminatory results, whether intentional or not, are unacceptable and are unlawful in New York.” The population map in step 4 is not a courtesy either.

In the municipal system the convenient proxies are the same shape. Repair-ticket closure time is cost, not severity. A confirmed fault is only ever confirmed where a crew was sent. Districts with fewer sensors and fewer reports generate fewer labelled incidents. The proxy is thinnest exactly where step 4 needs it thickest. Write down, before fitting anything, which quantity you would measure with an unlimited budget. Then write down how far each recorded field sits from it.

FigureProcess · 5 steps
  1. 1. Define the incident

    Group duplicate reports and sensor alerts into one operational event.

  2. 2. Set the cutoff

    Record what information existed when triage was performed.

  3. 3. Construct outcomes

    Use technician findings, repair logs, severity adjudication, and delayed resolution.

  4. 4. Map populations

    Separate asset types, districts, sensor coverage, weather, and reporting channels.

  5. 5. Record uncertainty

    Mark incomplete, disputed, and never-observed outcomes explicitly.

Comparison

Simulate the transfers the city will face

Use several protected evaluations rather than one random split. The transfer that breaks a triage model is usually a site, not a season.

The Epic Sepsis Model has been evaluated at two independent sites, and the two results are not the same result. At Michigan Medicine, Wong and colleagues reported an area under the curve of 0.63 (95% CI 0.62-0.64) across 27,697 patients and 38,455 hospitalizations. A second team ran the model across 145,885 encounters in two county emergency departments and published in 2024: sensitivity 14.7%, specificity 95.3%, positive predictive value 7.6%. Their conclusion was that it “provides suboptimal diagnostic characteristics for undifferentiated patients in a county ED”. Same model, a different health system. A district holdout is that experiment run before deployment instead of after it.

The four evaluations below each simulate a different transfer. None of them substitutes for another.

FigureComparison · 4 columns

Future-period holdout

Tests later weather, policy, and infrastructure conditions.

  • Chronological features
  • Delayed-label maturity
  • Seasonal coverage
  • Relevant to immediate deployment

District holdout

Tests transfer to neighborhoods with different assets and reporting patterns.

  • Spatial dependence
  • Equity relevance
  • Fewer independent units
  • May reveal sensor gaps

Asset-family holdout

Tests new equipment or infrastructure categories.

  • Cold-start behavior
  • Different failure modes
  • Supports scope restrictions
  • Can be very difficult

Prospective shadow trial

Runs recommendations without dispatch authority.

  • Measures live data quality
  • Tests latency and queues
  • Avoids immediate side effects
  • Still lacks intervention outcomes

Example

Build a metric portfolio around the workflow

The scorecard includes primary outcomes, guardrails, and diagnostics. Precommit all three, because a deployed triage system will otherwise be reported on whichever one looks best.

The Epic Sepsis Model was already widely implemented when Michigan Medicine validated it from outside. The validation ran to 27,697 patients and 38,455 hospitalizations, between 6 December 2018 and 20 October 2019. Wong and colleagues published it in JAMA Internal Medicine in 2021. The hospitalization-level area under the curve was 0.63 (95% CI 0.62-0.64). The operating point mattered more than the curve: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” Against timely clinical practice, the alerts added 183 sepsis patients, 7%.

That is a full portfolio, reported at once, years into deployment. A discrimination number. A miss rate at the threshold actually used. An alert volume. A slice of clinicians' attention consumed. A marginal benefit over the existing workflow. Every one of the seven lines below exists so the city has those numbers before the crews do, not after.

  • Primary ranking metric: recall of critical incidents within the top 50 cases per shift. The Epic Sepsis Model was never held to that quantity before rollout, and at its own alert threshold it missed 1,709 of 2,552 sepsis cases, 67%.
  • Calibration: reliability of severe-risk probabilities in operational score bands, reported band by band. One summary figure hides the operating point, as the 0.63 (95% CI 0.62-0.64) did.
  • Guardrail: false dispatches per crew-hour and repeated burden per address. It is the municipal counterpart of alerts on 18% of 38,455 hospitalizations, which the authors call “a large burden of alert fatigue”.
  • Summary checks: claim-level evidence support, numerical fidelity, and missing-information flags. Score them against the technician note, not against the draft's own fluency.
  • Equity slices: coverage, recall, delay, and abstention by district and reporting channel. The referral share moving from 17.7% to 46.5% is the warning of what a slice table catches and a global metric never does.
  • Robustness: performance under storms, missing sensors, low-quality photos, and communication outages. Measure each as a named condition rather than averaging it into the whole.
  • Outcome: time to safe resolution, secondary damage, crew utilization, and complaint or appeal rates, always as a margin over current practice. The sepsis model's honest number on that line was 183 patients, 7%.

Key idea

Precommit the release table — regulators now require it in writing

The city should not negotiate success after viewing the results. Define minimum recall, calibration tolerance, crew-burden ceiling, district slice floors, latency, and uncertainty rules in advance.

In one regulated domain this is no longer only good practice. On 4 December 2024 the FDA announced its final guidance on predetermined change control plans for AI-enabled device software. The mechanism is precommitment: “This guidance recommends that a PCCP describe the planned AI-DSF modifications, the associated methodology to develop, validate, and implement those modifications, and an assessment of the impact of those modifications.” The plan is submitted and reviewed as part of the marketing submission. Where the agency authorises it, the modifications it describes can then be implemented without an additional marketing submission. Writing the table down in advance is what buys the freedom to change the model later.

Note what the document is and is not. It is issued under the agency's good guidance practices, and it says of itself that it “does not establish any rights for any person and is not binding on FDA or the public”. A revised final version is dated August 2025. Even non-binding, it describes the artefact your capstone needs: which changes are foreseen, how each will be validated, and what impact each is expected to have.

A limited launch may be allowed when global requirements pass but one low-support slice remains uncertain. The scope must then exclude that slice, and targeted evidence must be collected.

Scope restriction is safer than pretending weak evidence is a pass.

Evaluate the human–AI workflow

Technicians and dispatchers should be tested with and without the system on realistic cases. Measure decision quality, time, override reasons, trust calibration, fatigue, and whether summaries anchor reviewers on incorrect assumptions. Randomize presentation order, blind the source of drafts where feasible, and retain expert adjudication for safety-critical disagreements.

Design that study as if the pairing might hurt, because the pooled evidence says it often does. The experiments have now been added up. Vaccaro and colleagues published, in Nature Human Behaviour in 2024, “a preregistered systematic review and meta-analysis of 106 experimental studies reporting 370 effect sizes”. MIT Sloan's news office counts the same base independently: “370 results...from 106 different experiments published in relevant academic journals and conference proceedings between January 2020 and June 2023”. The headline finding is not the one the phrase “human in the loop” implies: “First, we found that, on average, human-AI combinations performed significantly worse than the best of humans or AI alone.”

The losses were concentrated in decision-making tasks and the gains in content-creation tasks. The municipal system is both at once. Ranking and dispatch are decisions. Summary drafting is content. So the human study cannot be run once for the product. It has to be run per workflow step. The dispatch step is the one where the base rate says the pairing is most likely to lose to whichever half is better alone.

DECIDE-AI covers the earlier stage, when a system first meets live cases at small scale. It addresses “an AI system’s actual clinical performance at small scale”, and it “comprises 17 AI-specific reporting items (made of 28 subitems) and ten generic reporting items”. It exists because “the reporting of these early studies remains inadequate”. Staging is an obligation there, and reporting each stage is part of the obligation.

Figure

What two clinical guidelines require an AI study to report, counted — and the number of people who agreed each list.

Across 370 effect sizes, the pairing did worse on average than the better of its two halves.

Visual

Plan a staged deployment

Increase consequence only after the prior stage supplies adequate evidence. Design each stage so the stage below it survives when the stage above is withdrawn.

That withdrawal is not hypothetical. On 24 October 2023 the California DMV suspended Cruise LLC's autonomous-vehicle deployment and driverless testing permits. It cited the regulation covering vehicles “not safe for the public's operation”, and a second one on misrepresentation of safety information. The press release states the action and its limit in the same breath: “The California DMV today notified Cruise that the department is suspending Cruise’s autonomous vehicle deployment and driverless testing permits, effective immediately. The DMV has provided Cruise with the steps needed to apply to reinstate its suspended permits, which the DMV will not approve until the company has fulfilled the requirements to the department’s satisfaction. This decision does not impact the company’s permit for testing with a safety driver.”

Read that as a stage table collapsed in one direction. The highest-consequence stages — driverless deployment and driverless testing — ended overnight. The supervised stage, the one with a human able to take over, was left standing. That is the scope restriction the release table above is meant to make available in advance. Here a regulator executed it instead of the operator.

A second regulator was running its own proceeding over the same 2 October 2023 incident. On 1 December 2023 the California Public Utilities Commission issued an Order to Show Cause ruling concerning Cruise's interactions with the Commission after that incident. Stage 5 of any plan below therefore has two exits, not one. The evidence can fail. The account you gave of the evidence can fail.

FigureTimeline · 5 stops
  1. Stage 1: Offline replay

    Evaluate historical incidents with protected temporal and district splits.

  2. Stage 2: Shadow operation

    Run live ranking, summaries, and tools without changing dispatch.

  3. Stage 3: Assisted pilot

    Show recommendations to selected teams with mandatory confirmation.

  4. Stage 4: Controlled rollout

    Randomize eligible shifts or districts where interference is manageable.

  5. Stage 5: Monitored service

    Expand scope with drift, delayed outcomes, incident response, and rollback.

Steps

Design the live evidence loop

Every release assumption becomes a monitoring target, expressed as a number with a threshold and a named owner.

The assumptions that go stale first are the operating-point ones. An alert rate is a monitoring target because it drifts. Alerts reached 18% of 38,455 hospitalizations at that threshold in that period, and nothing in the model guarantees the same share next quarter. So is the miss rate, which is only visible once outcomes backfill. So is the slice table, which is where a cost-shaped proxy label silently reappears. Each of the five steps below turns one release-table row into a standing measurement.

FigureProcess · 5 steps
  1. 1. Monitor contracts

    Check feeds, timestamps, identifiers, duplicates, images, and label maturity.

  2. 2. Monitor decisions

    Track score distributions, thresholds, queue sizes, abstention, and overrides.

  3. 3. Backfill outcomes

    Update performance as repairs and investigations resolve.

  4. 4. Review slices

    Watch districts, asset types, weather, channels, and repeated addresses.

  5. 5. Trigger response

    Rollback, narrow scope, retrain, recalibrate, or collect evidence through named runbooks.

Example

The risk register

The final design memo must list unresolved risks and controls. A register entry is worth more when the risk carries a measured magnitude instead of an adjective.

Automation bias has one. Twenty-seven radiologists read 50 mammograms with a purported AI system. The first ten were a training set, on which the system's BI-RADS suggestions were correct. In the remaining 40, an incorrect BI-RADS category was suggested for 12. The share of mammograms rated correctly collapsed on the sabotaged cases. Inexperienced readers fell from 79.7% (SD 11.7) to 19.8% (SD 14.0). Moderately experienced readers fell from 81.3% (SD 10.1) to 24.8% (SD 11.6). Very experienced readers fell from 82.3% (SD 4.2) to 45.5% (SD 9.1), all P < .01. Dratsch and colleagues reported the experiment in Radiology in 2023. Experience halved the damage and did not prevent it: “The results show that inexperienced, moderately experienced, and very experienced radiologists reading mammograms are prone to automation bias when being supported by an AI-based system.”

The experiment also shows what a register entry should contain. The mechanism. The population it was measured in. The size of the effect. The condition under which it appears. Here that condition is a wrong suggestion presented with the same confidence as a right one.

  • Selective labels: serious cases receive richer investigation than routine cases, biasing outcome visibility. As in the cost-proxy health algorithm, the bias lands in the label rather than in the model.
  • Feedback loop: faster dispatch can reduce observed severity and change future labels. The outcome series then stops being comparable to the one the release table was written against.
  • Geographic inequity: areas with fewer sensors may be deferred more often or escalated more slowly. The referral share moving from 17.7% to 46.5% under a corrected target is the scale such a gap can reach.
  • Automation bias: staff may trust polished summaries despite unsupported details. It was measured at 79.7% down to 19.8% correct for the least experienced readers, and still 82.3% down to 45.5% for the most experienced.
  • Capacity shock: a storm can change prevalence and overwhelm the top-k queue, converting a tuned threshold into an alert volume nobody agreed to.
  • Security: malicious reports or prompt injection can manipulate summaries and tool recommendations. The automation-bias figures describe what happens downstream of a confident wrong suggestion, however it was produced.

Analogy

An air-traffic replacement that kept the old system running: FAA ERAM

The FAA's En Route Automation Modernization system replaced a 40-year-old Host Computer System at the agency's 20 Air Route Traffic Control Centers. It was planned to be operational at all 20 centers in December 2010. The Department of Transportation's Office of Inspector General records what happened instead: “The Agency planned to make ERAM operational at its 20 Centers in December 2010, but due to a series of software-related problems, ERAM’s implementation was delayed. FAA declared ERAM fully operational in March 2015”. The FAA's own programme page now states that “ERAM is utilized by air traffic controllers at all 20 en route centers in the Continental United States”, with expansion to Honolulu and Anchorage starting in 2027.

What makes this an evaluation programme rather than a late project is the route back. Throughout, the legacy Host's backup was kept alive: “FAA had retained the Host’s backup system, known as Enhanced Backup Surveillance (EBUS) system, until ERAM was more mature.” It was removed only after a two-year safety evaluation. The fallback was retired on evidence, on its own schedule, not on the new system's launch date.

What the staging buys is not confidence. It is the ability to stop. Every centre cut over was a centre the old system could still have carried. An evidence program with no route back has staged nothing at all. December 2010 to March 2015 is what the delay cost. A lost fallback would have cost something that does not appear in a schedule.

A credible evaluation program controls both uncertainty and deployment consequence.

The final evaluation dossier

The student delivers an intended-use statement, data and label contract, split design, metric portfolio, uncertainty plan, slice table, and robustness matrix. The package also includes a human-study protocol, release criteria, staged rollout, monitoring runbook, and risk register.

The conclusion must be bounded: full release, restricted release, more evidence, redesign, or rejection. “Promising” is not a disposition. The Epic Sepsis Model was in wide use before anyone published that it missed 67% of sepsis cases at its alert threshold. The health-risk algorithm was managing populations before anyone published that its label was cost. In both cases the dossier existed. It was simply written afterwards, by other people.

Clinical medicine has already written down what the report of an AI trial must contain. The CONSORT-AI extension, published in Nature Medicine in 2020, “includes 14 new items that were considered sufficiently important for AI interventions that they should be routinely reported in addition to the core CONSORT 2010 items”. Those items were settled through “a two-stage Delphi survey (103 stakeholders)” and “a two-day consensus meeting (31 stakeholders)”. What they ask for maps onto this capstone almost line by line: “the setting in which the AI intervention is integrated, the handling of inputs and outputs of the AI intervention, the human–AI interaction and provision of an analysis of error cases”. A dossier is the deliverable there too.

The capstone succeeds when the evidence leads to a defensible action, including the decision not to deploy.

Key takeaways