Evaluation
Thresholds, Costs, Utilities, and Capacity
Design decision thresholds using error consequences, prevalence, capacity, abstention, and multi-stage workflows rather than default cutoffs.
By the end you can
- Separate ranking scores, calibrated probabilities, and final decisions
- Choose thresholds from costs, utilities, and capacity constraints
- Design tiered decisions and abstention rather than one binary cutoff
- Re-evaluate thresholds when prevalence, workflow, or resources change
A model score is not yet an action
A classifier may output 0.63. That number does not say whether to block, review, defer, or ignore. The action comes from a policy that combines the score with costs, capacity, eligibility, and sometimes other evidence. Treating 0.5 as a universal default confuses a mathematical convention with an operational decision.
Medicine formalized this in 2006. Decision curve analysis starts from the patient rather than from the classifier. Vickers and Elkin begin by assuming that the threshold probability of a disease or event at which a patient would opt for treatment “is informative of how the patient weighs the relative harms of a false-positive and a false-negative prediction”. That relationship “is then used to derive the net benefit of the model across different threshold probabilities”. “Plotting net benefit against threshold probability yields” what the authors call the decision curve. The apparatus is modest. The method “requires only the data set on which the models are tested”. What it returns is not one cutoff. It is a curve across all of them.
The model orders or scores cases; the policy decides what happens next.
Visual
From output to intervention
Several transformations can sit between a model and an action. One of them has been published in full by a regulator. Apple's Irregular Rhythm Notification Feature received a De Novo classification, DEN180042, on 8 August 2018, and the confirmation layer is written into the FDA decision summary: “If a sufficient number of tachograms are retrieved and classified to meet the notification threshold (5 of 6 sequential tachograms classified as irregular within a 48-hour period), a notification indicating that the heart rhythm has shown signs of AF will be displayed to the user.”
Read that as the diagram above. The model output is a per-tachogram classification. The decision policy is a rule over six of them inside a 48-hour window. The operational action is one notification to one person. The Apple Heart Study followed 419,297 participants over a median 117 days of monitoring. That policy notified 2,161 of them, 0.52%, with a positive predictive value of 0.84 (95% CI 0.76 to 0.92) for atrial fibrillation on ECG concurrent with a subsequent notification. Nothing in the classifier fixes the 0.52%. The layer above it does.
- 01
Model output
Logit, margin, distance, probability estimate, or ranking score.
- 02
Calibration or mapping
Optional transformation that relates scores to empirical risk.
- 03
Decision policy
Thresholds, tiers, abstention, queue limits, and eligibility rules.
- 04
Operational action
Review, block, route, monitor, request evidence, or take no action.
Comparison
Expected cost depends on more than two constants
Cost matrices are useful. Real consequences are richer than two constants, and one recommended cutoff can turn an uncommon event into a permanent alarm.
Michigan Medicine's own hospital operations committee set its sepsis alert at a score of 6, inside the 5-8 range the model's developer recommends. Wong and colleagues then validated the Epic Sepsis Model against what that produced: 38,455 hospitalizations among 27,697 patients, containing 2,552 sepsis cases — 6.6% of hospitalizations. At that threshold the hospitalization-level AUC was 0.63 (95% CI 0.62-0.64). Sensitivity was 33%. Positive predictive value was 12%. The abstract puts the operational consequence plainly: “The ESM did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” A policy that fires on 18% of every admission is a staffing decision wearing a statistical costume. Ostermayer and colleagues repeated the exercise at Harris Health, across 145,885 encounters at the same score threshold of 6. Positive predictive value there was 7.6%.
One level up, a national guideline body has documented the same problem in its own rules. NICE sets urgent cancer referral at a 3% positive predictive value. An editorial in the British Journal of General Practice records where that number came from: “This was a reduction from the previous 5% threshold and was based on clinical consensus balancing diagnosing more cancers in a timely manner against overburdening the health service and potential implications for individuals (for example, unnecessary anxiety).”
The same editorial shows what the cutoff means once the population under it moves. The PPV of CA125 at the usual 35 U/mL threshold is 9%. The individualised risk of a woman whose CA125 is 35 U/mL is about 1%. A 45-year-old with CA125 38 U/mL has an estimated 0.6% ovarian cancer risk. A 70-year-old with the identical value has 3.1%. One test, one number, a five-fold difference in what it means.
Simple cost threshold
Assign one cost to each error type and choose the lower expected cost.
- Transparent starting point
- Assumes stable probabilities
- Often ignores capacity
- May omit repeated burden
Capacity-constrained policy
Select the highest-risk cases within a fixed review budget.
- Matches limited queues
- Threshold changes with score distribution
- Needs ranking quality
- Can leave moderate risks untouched
Tiered intervention
Use different actions for high, medium, and low risk.
- Supports escalation
- Can request more evidence
- Requires action-specific evaluation
- Adds policy complexity
Example
The same reading task under two screening policies
Two population breast-screening programmes have measured what changes when a model score routes cases through a reading workflow rather than replacing the reader. The interesting differences are not accuracy differences. They are reviewer-hour differences.
- MASAI, the randomised trial: 80,033 Swedish women, AI-supported screen reading against standard double reading. Lang and colleagues published it in The Lancet Oncology in 2023.
- Detection under the two policies: 244 cancers found with AI-supported reading against 203 with standard double reading, or 6.1 against 5.1 per 1,000.
- Capacity is the variable that actually moved. The same trial cut screen-reading workload by 44.3% — a change in who reads what, not a change in what the model computes.
- PRAIM, the nationwide implementation in Germany: Eisemann and colleagues screened 463,094 women, 260,739 of them with AI support. Their report in Nature Medicine: “Radiologists in the AI-supported screening group achieved a breast cancer detection rate of 6.7 per 1,000, which was 17.6% (95% confidence interval: +5.7%, +30.8%) higher than and statistically superior to the rate (5.7 per 1,000) achieved in the control group.”
- Recall rate is the guardrail on the routing tier, and it was noninferior at 37.4 against 38.3 per 1,000. Watch that number. A routing policy is judged by the actions it generates, not only by the cancers it finds.
Analogy
A thermostat with several modes
A thermostat can heat, ventilate, warn, or remain idle, and it chooses between them from temperature and occupancy together. One sensor reading never determines the action on its own. The control policy carries goals and constraints the sensor knows nothing about.
Temperature at least means the same thing on Tuesday as it did on Monday. A model score does not, and neither does a laboratory value read against a fixed bar. CA125 at 38 U/mL carries an estimated 0.6% ovarian cancer risk for a 45-year-old and 3.1% for a 70-year-old. The threshold belongs to the surrounding system. It has to be revisited whenever the score, or the population under it, shifts.
The same signal can support different actions under different policies.
Key idea
Thresholds are not portable by default
Changes in prevalence, calibration, reviewer capacity, error cost, or score distribution can make yesterday's threshold inappropriate. A threshold tied to the top 1,000 cases behaves differently from a threshold tied to a fixed probability.
Version thresholds and evaluate them as policy artifacts. Monitor the input population and the action volume after release. Two US bodies did exactly that with the same eligibility rule, eleven months apart. On 9 March 2021 the U.S. Preventive Services Task Force issued its Grade B recommendation on lung cancer screening: “The USPSTF recommends annual screening for lung cancer with low-dose computed tomography (LDCT) in adults aged 50 to 80 years who have a 20 pack-year smoking history and currently smoke or have quit within the past 15 years.” The starting age had dropped from 55 to 50, the smoking history from 30 to 20 pack-years. No model was retrained. What moved was the eligible population: modelling put the relative increase at 87% overall — 78% in non-Hispanic White adults, 107% in non-Hispanic Black adults, 112% in Hispanic adults. Then CMS rewrote national coverage determination CAG-00439R on 10 February 2022 to cover ages 50 to 77 with at least 20 pack-years. The payer's version of the threshold and the task force's version do not share an upper bound.
The problem underneath all of this got its name in 2000. Provost and Fawcett put it in one sentence: “In real-world environments it usually is difficult to specify target operating conditions precisely, for example, target misclassification costs.” Their answer was not a sharper number. It was the ROC convex hull, a comparison method “that is robust to imprecise class distributions and misclassification costs”, from which “it is possible to build a hybrid classifier that will perform at least as well as the best available classifier for any target conditions”. The question is not which threshold is right today. It is which set of thresholds survives being wrong.
A fixed number is not a fixed operating condition.
Position
The threshold is somebody's decision about whose error costs more
Somebody chose the cutoff. When nobody remembers choosing it, the value is 0.5, a mathematical convention rather than anything the model possesses. The convention hides the question the line actually answers: which error is worse, and for whom.
Sometimes the choosers write it down. NG12, the NICE guideline on the recognition and referral of suspected cancer, sets urgent referral at a 3% positive predictive value of cancer, and the guideline text names the body that agreed it: “The GDG agreed to use a 3% PPV threshold value to underpin the recommendations for suspected cancer pathway referrals and urgent direct access investigations.” No model produced that number. It carries explicit exceptions — for children and young people, and for tests routinely available in primary care — and each exception is itself a decision about who is worth investigating at a lower bar. Decision curve analysis assumes precisely this. It treats the threshold probability at which a patient would opt for treatment as something that “is informative of how the patient weighs the relative harms of a false-positive and a false-negative prediction”. Whoever sets the cutoff is doing that weighing on the patient's behalf, whether or not anyone says so.
The cutoff does not have to move for the people on the wrong side of it to change. US kidney transplant listing turns on an eGFR of 20 mL/min per 1.73 m2. Nobody changed 20. What changed was the score feeding it. The OPTN required race-neutral eGFR, effective 5 January 2023, with programmes given until 3 January 2024 to submit. Schold and colleagues report the result: “Overall, 32% (14,419/44,912) of Black candidate listings received an eGFR modification of waiting time priority.” The median increase was 610 qualifying priority days (IQR 330-1049), and those listings went on to a deceased-donor transplant rate 2.85 times higher (95% CI 2.70 to 3.02). A fixed threshold reassigned 14,419 listings by standing still while the measurement underneath it was rewritten.
The remedy is not a better default. In 2000 Provost and Fawcett wrote that in real-world environments it usually is difficult to specify target operating conditions precisely, target misclassification costs among them. Their answer was the ROC convex hull, a comparison method “that is robust to imprecise class distributions and misclassification costs”, rather than a sharper single number. The useful question to put to a deployed system is not what its threshold is. It is who set it, against which costs, and on which definition of the score. And what happens to the people on the wrong side of it when prevalence or reviewer capacity shifts underneath.
An unexamined default is still a decision. It is only an anonymous one.
Borderline cases deserve a policy
When uncertainty is high or consequences are severe, abstention can be better than forced classification. The system can request more information, route to a specialist, or delay action. The Irregular Rhythm Notification Feature is an abstention policy read from the other end. Five of six sequential tachograms have to agree within 48 hours before anything is said to the user. Everything short of that is a decision not to decide yet.
Abstention must be evaluated for coverage, residual error, subgroup burden, delay, and fallback quality. It is not a free safety feature. The 67% of septic patients the Epic Sepsis Model did not identify at a score threshold of 6 were, from the alerting system's point of view, cases it declined to flag. The fallback pathway is what decided their outcome.
Declining to decide moves risk into the fallback pathway.
Steps
Choose and document the operating policy
The threshold should be reproducible and reviewable. The FDA decision summary for the Apple notification feature is what the last step of this process looks like when it is done properly. The notification rule — 5 of 6 sequential tachograms classified as irregular within a 48-hour period — is written where a reader outside the company can find it, read it, and check the action volume it produced. NG12 is the same discipline applied to a committee's judgement rather than a device's firmware. The 3% bar, the body that agreed it, and the cases exempted from it are all on the record. Neither document tells you the threshold is correct. Both make it possible to argue about. That is the property a versioned policy is supposed to have.
1. Define actions
List every possible intervention, including abstention and no action.
2. Quantify constraints
Record capacity, latency, legal limits, and severity-dependent costs.
3. Evaluate curves
Measure outcomes across thresholds and candidate tiers on protected data.
4. Stress assumptions
Vary prevalence, calibration, costs, and available reviewer hours.
5. Version the policy
Store thresholds, tie-breaking, overrides, monitoring, and rollback conditions.
Key takeaways
- Model outputs become actions only through an explicit decision policy. Five of six sequential tachograms inside a 48-hour window is the policy, and it produced a 0.52% notification rate across 419,297 participants.
- Threshold choice depends on prevalence, calibration, error consequences, and operational capacity. A 6.6%-prevalence sepsis problem became alerts on 18% of all hospitalizations, at 12% positive predictive value, at a score threshold of 6.
- Fixed-capacity queues create moving score cutoffs and put the weight on ranking quality. MASAI cut screen-reading workload by 44.3% across 80,033 women without changing what counted as a cancer.
- Tiered actions can be more appropriate than one binary threshold, and the tiers are policy. NG12's 3% bar carries agreed exceptions for children and young people and for tests routinely available in primary care.
- Abstention transfers risk to a fallback and requires its own evaluation. The 1,709 septic patients missed at a threshold of 6 were handled by whatever the fallback pathway was.
- Thresholds should be versioned, monitored, stress-tested, and revised when operating conditions change. The USPSTF moved 55 to 50 and 30 to 20 pack-years on 9 March 2021, an 87% relative increase in the eligible population, and CMS followed with CAG-00439R on 10 February 2022.