Evaluation
Evaluation as Decision Evidence
Build a system-level view of evaluation that connects intended use, evidence, uncertainty, risk, and release decisions.
By the end you can
- Describe evaluation as evidence for a decision rather than a search for one impressive score
- Distinguish model quality, system quality, and operational outcome quality
- Identify the claims that an evaluation can and cannot support
- Write a decision-focused evaluation brief before selecting metrics
Four correct numbers, one model, one dataset
On the same 38,455 hospitalizations at Michigan Medicine, the Epic Sepsis Model scored an area under the curve of 0.63, a sensitivity of 33%, a specificity of 83% and a negative predictive value of 95%. Every one of those numbers is correct. All four were computed on the same patients, at the same alert threshold of 6 the hospital was actually using. A buyer handed the last of them learns nothing about the two thirds of sepsis patients the model failed to identify.
Evaluation begins by naming the decision the evidence has to support. A metric becomes useful only once its population, unit, time window, workflow and acceptable failure modes are explicit. As often as not, it becomes useful only once somebody outside the vendor has computed it.
A valid calculation can still be irrelevant to the decision.
Case
An AUC of 0.63 that missed 1,709 of 2,552 sepsis patients
The Epic Sepsis Model was already widely implemented when somebody outside the vendor measured it. The external validation ran at Michigan Medicine and appeared in JAMA Internal Medicine on 21 June 2021. It covered 27,697 patients and 38,455 hospitalizations, and reported a hospitalization-level area under the curve of 0.63 (95% CI 0.62–0.64).
At the alert threshold of 6 the hospital used, sensitivity was 33%, specificity 83%, positive predictive value 12% and negative predictive value 95%. The model missed 1,709 of the 2,552 patients who developed sepsis, 67% of them. It alerted on 6,971 hospitalizations, 18% of the total. Wong and colleagues put both halves in one sentence: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.”
None of those figures is a mistake. They are answers to four different questions. And the paper's title — External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients — records which question the deployment had never asked.
Comparison
Four claims that require different evidence
Teams often compress several claims into the sentence “the model performs well.”
Google's retinopathy screening model went to eleven clinics in Thailand, and there the claims came apart. In laboratory evaluation it had been reported at over 90% accuracy and described as specialist-level. Beede and colleagues followed it into the clinics in 2020. More than a fifth of the images the nurses captured were refused by the system's own quality threshold, and slow internet connections added further delay. MIT Technology Review put the clinic conditions in one sentence: “With nurses scanning dozens of patients an hour and often taking the photos in poor lighting conditions, more than a fifth of the images were rejected.”
The predictive claim travelled to Thailand intact. The system claim did not. No amount of additional accuracy would have recovered a patient whose photograph the model declined to read.
Predictive quality
Outputs agree with reference outcomes on a defined sample.
- Requires a trustworthy reference
- Depends on the sampling scheme
- May change with threshold
- Does not prove workflow value
Decision quality
Actions based on outputs improve choices under stated costs.
- Requires a decision rule
- Needs error consequences
- Includes abstention or fallback
- May favor a simpler model
System quality
The complete service behaves reliably under operational constraints.
- Includes latency and availability
- Tests integration failures
- Covers monitoring and rollback
- Depends on human interaction
Outcome quality
Real-world results improve for affected people or processes.
- Often needs online or prospective evidence
- Can involve delayed outcomes
- Must examine harms and tradeoffs
- May differ from offline metrics
Visual
The evidence chain from purpose to action
Each link narrows what the final numbers mean, and every failure in this lesson is a link that nobody wrote down. The Epic Sepsis Model had an operating point — a score of 6 — before it had a local validation. Ofqual's standardisation model had a defensible claim about centres and an indefensible one about students. Google Flu Trends got its first post-release check when Lazer and colleagues ran one, eight years in.
1. Intended use
Specify who uses the system, for what decision, and under which constraints.
2. Evaluation claim
State exactly what improvement or safety property is being tested.
3. Evidence design
Choose units, samples, references, metrics, and uncertainty estimates.
4. Decision rule
Define pass, fail, revise, abstain, or collect more evidence.
5. Post-release check
Monitor whether assumptions and outcomes remain credible after deployment.
Example
One model, four verdicts, and the queue it created
Take the Michigan Medicine figures and put each of the four claims to them in turn. The dataset never changes. The verdict does.
- Predictive quality: a hospitalization-level area under the curve of 0.63, with a 95% confidence interval of 0.62–0.64. That is a statement about how the model orders patients on those 38,455 hospitalizations. It says nothing else.
- Decision quality: at the threshold of 6 the hospital actually used, sensitivity was 33% and positive predictive value 12%. Most septic patients were not flagged, and most flagged hospitalizations were not septic. Clinicians met the decision rule, not the ranking.
- System quality: 6,971 alerts across 38,455 hospitalizations, 18% of them. That is the queue the model created for the people around it, and Wong and colleagues named the result “a large burden of alert fatigue”.
- Outcome quality: specificity 83% and negative predictive value 95% are the two most flattering numbers in the paper. Neither is evidence that a single patient was treated sooner. The validation measured the model. It did not measure the care.
Analogy
A courtroom with several standards of proof
In a courtroom, each claim carries its own burden of proof. A witness can establish that an event occurred without proving motive, liability, or the appropriate remedy.
Evaluation fixes no burdens in advance. Nobody rules a result inadmissible for the claim it is being asked to carry, so matching the evidence to the claim is left to whoever is reading the number. An area under the curve of 0.63 is testimony about ranking. For as long as the Epic Sepsis Model was widely implemented, it was read as testimony about patient safety, and nothing in the courtroom stopped it.
Do not let one piece of evidence impersonate an entire case.
Key idea
The metric-first trap, and the benchmark that moved 40 percentiles
Teams often begin by asking whether to use accuracy, F1, or AUC. That order is backwards. Metric choice depends on the decision, the prevalence, the capacity, the costs and the quality of the reference.
The second trap is accepting a benchmark as a product verdict. OpenAI's GPT-4 technical report opens with one, in March 2023: “GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers.”
Eric Martínez re-examined that result in 2024. He left the raw score untouched and changed only the population it was compared against. The 90th-percentile figure came from February administrations, which are skewed toward repeat takers. Against a July cohort, the same performance sits below the 69th percentile overall and at the 48th on essays. Against first-time takers, the 62nd percentile. Against those who actually passed, the 48th percentile overall and the 15th on essays.
Not one token of model output changed between the top 10% and the 15th percentile. The comparison population did. That is the opening point of this lesson arriving as a headline.
Choose the claim first, then design the evidence.
Steps
Write a decision-focused evaluation brief
Use this before calculating any headline metric. Three regulators already require most of it in writing. On 27 October 2021 the U.S. FDA, Health Canada and the UK MHRA jointly issued ten guiding principles for Good Machine Learning Practice for Medical Device Development. Principle 7 refuses to accept a model score on its own: “Where the model has a 'human in the loop,' human factors considerations and the human interpretability of the model outputs are addressed with emphasis on the performance of the Human-AI team, rather than just the performance of the model in isolation.” Principle 8 requires test plans covering the intended patient population, important subgroups, the clinical environment, measurement inputs and potential confounders. Steps 2, 3 and 4 below are not this lesson's advice. They are a regulatory expectation with a date on it.
Step 2 is the one that gets skipped, and England's summer 2020 A-levels show the price. With exams cancelled, Ofqual awarded grades from its Direct Centre Performance standardisation model. That model was built to reproduce a centre's grade distribution, and it was evaluated at that unit. The decision it was actually making was about an individual student's university place. On results day, 13 August 2020, 39.1% of entries were downgraded from the centre assessment grade and only 2.2% were raised. Ofqual's own evaluation records 39% of entries lowered and just over 2% raised: “in most instances where there were differences, teachers' assessments of grades were higher than the results produced by the standardisation method (39% of entries)”. The calculated grades were abandoned four days later. A model can be right about the unit it was measured on and catastrophic at the unit it decides.
1. Name the decision
Write the release, rollback, routing, or research choice the evaluation will inform.
2. Define the unit
Specify whether evidence concerns users, sessions, patients, devices, queries, or events.
3. Describe the population
State geography, time period, inclusion rules, and important subgroups.
4. List failure costs
Record false positives, false negatives, delay, overload, and silent failure.
5. Precommit thresholds
Set acceptance criteria and what happens when evidence is inconclusive.
A bounded statement is now a legal obligation, not a courtesy
The final report should say what was tested, on whom, under which protocol, with what uncertainty, and for which decision. It should also state what remains unknown.
That discipline stopped being a matter of professional taste on 13 June 2024, the day the EU adopted the Artificial Intelligence Act. Article 15(3) of Regulation (EU) 2024/1689 provides: “The levels of accuracy and the relevant accuracy metrics of high-risk AI systems shall be declared in the accompanying instructions of use.” Both halves are named. Not the level alone, which is the marketing number. Not the metric alone, which is the technical alibi. And they travel with the product, to the person deploying it. Article 15(2) then tasks the Commission with encouraging the development of benchmarks and measurement methodologies for those levels — a legislature conceding that the metric worth declaring is not yet settled.
That requirement prevents a narrow score from becoming a broad marketing claim, and it makes later monitoring possible.
The most trustworthy conclusion is often narrower than the most exciting one.
Case
The algorithm that measured cost, reported health, and got a letter
A commercial risk algorithm ranked patients by predicted health cost rather than by illness. Obermeyer and colleagues took that substitution apart in Science, in the issue dated 25 October 2019. The system was “affecting millions of patients”; it “predicts health care costs rather than illness”; and at a given risk score, Black patients were “considerably sicker” than White patients. Remedying the disparity would raise the share of Black patients receiving additional help “from 17.7 to 46.5%”. The algorithm computed exactly what it had been asked to compute.
What turned that finding into evidence was the decision it fed. A joint letter dated the same day, 25 October 2019, went from the New York Department of Financial Services and Department of Health to UnitedHealth Group CEO David S. Wichmann about Optum's Impact Pro algorithm. It demanded that the company stop using the algorithm or demonstrate that it is not discriminatory: “These discriminatory results, whether intentional or not, are unacceptable and are unlawful in New York.” An evaluation is evidence for a decision, and the decision here belonged to somebody with the power to stop the deployment.
Position
Deployment at scale is not evidence that a model works
How many hospitals already run it is the first thing a buyer learns about a model, and the least informative thing about it. The Epic Sepsis Model was already widely implemented when the Michigan Medicine validation put numbers on it, across 27,697 patients and 38,455 hospitalizations: an area under the curve of 0.63 (95% CI 0.62–0.64), sensitivity 33% at the alert threshold of 6, and 1,709 of 2,552 sepsis patients not identified, 67% of them. Alerts fired on 6,971 hospitalizations, which Wong and colleagues called “a large burden of alert fatigue”. Being installed everywhere had predicted none of it.
The cost-proxy algorithm fails in the opposite direction, and the pair is worth reading together. Its arithmetic was never wrong. It ranked patients by predicted health cost, correctly, and that ranking decided who received additional help. Fixing the target would have raised the share of Black patients receiving that help “from 17.7 to 46.5%”. One system was weak at the thing it was there to do. The other did exactly what it had been built to do, and the wrong thing had been built. Counting installations separates neither case from a system that works, and in both the answer arrived when somebody outside the vendor measured.
Google Flu Trends was the exhibit for the opposite proposition: free, global, continuously running, refreshing faster than any agency could publish, and cited over and over as the exemplary big-data system. Then Lazer and colleagues measured it in Science on 14 March 2014: “GFT overestimated the prevalence of flu in the 2012–2013 season and overshot the actual level in 2011–2012 by more than 50%. From 21 August 2011 to 1 September 2013, GFT reported overly high flu prevalence 100 out of 108 weeks.” A model that did nothing but lag the CDC's own figures had a mean absolute error of 0.311. The system running live on the whole world's searches had 0.486. Scale and uptime were the reasons everyone trusted it, and they were measurements of nothing.
“Widely used” is a fact about procurement, not a measurement of performance.
Key takeaways
- Evaluation is a chain from intended use to evidence, decision rule and post-release check. The Epic Sepsis Model was widely implemented before anyone measured an area under the curve of 0.63 on 38,455 hospitalizations.
- Model quality, decision quality, system quality and outcome quality are four different claims. Google's retinopathy model was reported at over 90% accuracy in the lab and had more than a fifth of its images rejected across eleven clinics in Thailand.
- A metric means nothing without a defined unit, population, reference, threshold and time window. GPT-4's bar exam result moved from the top 10% to the 48th percentile overall on nothing but a change of comparison cohort.
- Operational capacity can reverse the verdict an offline score suggests: 6,971 alerts across 38,455 hospitalizations was “a large burden of alert fatigue”, whatever the ranking metric said.
- Acceptance criteria belong in writing before results are inspected. Principles 7 and 8 of the FDA, Health Canada and MHRA guiding principles of 27 October 2021 require the test plan to name the human-AI team, the population, the subgroups and the environment.
- A trustworthy conclusion states what the evidence supports and what it leaves unresolved. Article 15(3) of the EU's Artificial Intelligence Act turns that into a declaration of levels and metrics shipped with the product.