Evaluation
Metric Portfolios, Acceptance Criteria, and Release Gates
Combine primary metrics, guardrails, slices, uncertainty, and operational constraints into a release decision that resists cherry-picking.
By the end you can
- Build a metric portfolio around a primary decision claim
- Distinguish optimization, diagnostic, guardrail, and operational metrics
- Predefine acceptance criteria that include uncertainty and tradeoffs
- Resolve conflicting metrics without inventing a score after results are known
Example
A dashboard is not a decision rule
A single headline number would have passed the Epic Sepsis Model. Wong and colleagues measured it from outside. Across 38,455 hospitalizations at Michigan Medicine the model's area under the curve was 0.63 (95% CI 0.62–0.64), against the vendor's claimed 0.76–0.83. The portfolio failure is in the next sentence of the abstract: “The ESM also did not identify 1709 patients with sepsis (67%) despite generating alerts for an ESM score of 6 or higher for 6971 of all 38 455 hospitalized patients (18%), thus creating a large burden of alert fatigue.” That is Wong, Otles and colleagues, in JAMA Internal Medicine in 2021. Miss 67% of sepsis cases and alert on 18% of all hospitalizations at the same time, and you have a number needed to evaluate of 8. Eight alerts read for one case found.
A second, independent validation reached the same verdict on different patients. In two Houston county emergency departments, across 145,885 encounters, the model showed sensitivity of 14.7%, specificity of 95.3% and a positive predictive value of 7.6%. Ostermayer and colleagues concluded that it “provides suboptimal diagnostic characteristics”. Note which number is the high one. Specificity of 95.3% is the figure a dashboard leads with, and it is the figure that decides nothing here. Teams collect dozens of numbers and choose afterward which ones mattered. A portfolio decides first. It names each number's job before any of them are read.
- Primary metric: whether the deteriorating patients are actually identified. At Michigan Medicine, 67% of sepsis cases were missed. In the Houston replication, sensitivity was 14.7%. That is the quantity the system exists to deliver.
- Guardrail: alert burden per clinician. Alerts on 6971 of 38 455 hospitalized patients (18%), a number needed to evaluate of 8, a positive predictive value of 7.6% — an operating point clinicians cannot work at, whatever the discrimination.
- Slice criterion: hold performance to a floor in every population that matters. Two emergency departments in one county produced 14.7% sensitivity; a pooled validation reported an AUC of 0.63. An average across sites would have shown neither.
- Operational constraint: the interaction budget. At Bing, 100 milliseconds of delay was priced at 0.6% of revenue. Feasibility is a release criterion, not a footnote.
- Uncertainty rule: read the bound, not the point. The IDx-DR pivotal trial required a one-sided 97.5% bound to clear thresholds of 85.0% sensitivity and 82.5% specificity. An AUC of 0.63 (95% CI 0.62–0.64) is precise enough to convict. Many results are not.
Visual
Four roles for metrics
Assign the roles before the numbers arrive and no measurement has to compete for headline status. The sepsis validations supply one of each.
The primary outcome is the quantity attached to the decision: whether patients with sepsis are identified. 67% were not. A guardrail is a harm, cost or degradation that must not exceed its limit. Alerts on 18% of all hospitalizations, a number needed to evaluate of 8, a positive predictive value of 7.6% — that is what a guardrail looks like when nobody wrote one down. A diagnostic explains behaviour and directs improvement without deciding anything. An AUC of 0.63 (95% CI 0.62–0.64) against a claimed 0.76–0.83 tells you the model does not discriminate as advertised, but the release decision was already made by the two numbers above it. An operational measure decides feasibility: latency, capacity, availability, cost, reviewer workload. It is where the clinician's working day meets the model.
The same four figures, unlabelled and side by side, are a dashboard. The specificity of 95.3% is the one that gets quoted in the slide. Labelled by role, they are the beginning of a decision rule, and 95.3% is visibly answering a question nobody needed answered.
Primary outcome
The main quantity connected to the intended decision or user value.
Guardrail
A harm, cost, or degradation that must not exceed its limit.
Diagnostic
A measure used to explain behavior and guide improvement.
Operational
Latency, capacity, availability, cost, or workload needed for feasibility.
Comparison
Metric conflict is information
The price of leaving the analysis unfixed until the results are in has been measured twice, in two disciplines, with the same design. Hand many independent teams one dataset and one hypothesis. Count how far they diverge.
Twenty-nine teams — 61 analysts — were given identical soccer red-card data and a single primary research question. Silberzahn, Uhlmann and colleagues reported the spread in 2018: “The independent teams’ estimated effects for the primary research question ranged from 0.89 to 2.93 in OR units (1.0 indicates a null effect); no teams found a negative effect, 9 found no significant relationship, and 20 found a positive effect.” Twenty of twenty-nine — 69% — certified the effect. The remaining 9 found nothing to certify. Nobody cheated and nobody saw different data.
Botvinik-Nezer and colleagues repeated the experiment in neuroimaging: 70 teams, one fMRI dataset, the same 9 ex-ante hypotheses, 70 different analysis pipelines. From the abstract, in Nature: “The flexibility of analytical approaches is exemplified by the fact that no two teams chose identical workflows to analyse the data.” Not two workflows in agreement out of seventy.
That is what a post hoc score is worth. Inventing the combination rule after seeing which model wins does not merely invite cherry-picking, weaken reproducibility and hide tradeoffs. It places the conclusion inside a 69%/31% band that the analyst's own choices control.
A precommitted hierarchy removes that latitude: primary claim, guardrails, tie-breakers, written before evaluation. It exposes failure directly and still allows a documented exception. Where no hierarchy commands agreement, a Pareto analysis is the honest alternative — show every candidate not dominated across the chosen dimensions, avoid an arbitrary scalarization, and make the stakeholders choose in public. All three responses to conflict are available. Only the first is chosen after the answer is known.
Post hoc score invention
Combine metrics after seeing which model wins.
- Invites cherry-picking
- Weights lack prior meaning
- Hides tradeoffs
- Weakens reproducibility
Precommitted hierarchy
Define primary, guardrails, and tie-breakers before evaluation.
- Clarifies priorities
- Supports consistent release rules
- Exposes failure directly
- Still allows documented exceptions
Pareto analysis
Show models that are not dominated across key dimensions.
- Preserves tradeoffs
- Avoids arbitrary scalarization
- Requires stakeholder choice
- Useful for cost-quality frontiers
Steps
A release gate should be executable
There is a dated, published example of a gate written as a policy rather than presented as a value. It is worth following step by step.
1. State the primary claim. The IDx-DR pivotal trial named two: sensitivity and specificity for detecting more-than-mild diabetic retinopathy, in primary-care offices, on the patients the device was meant to serve.
2. Set minimum guardrails before enrolment. The FDA's own De Novo summary records the numbers as they stood before any data existed: “Performance thresholds were defined at 85.0% for sensitivity and 82.5% for specificity”. Not a target. Not an aspiration. A rejection region fixed in the protocol.
3. Add slice floors. An aggregate can be strong while a slice is catastrophic, and the gap has been measured across two orders of magnitude. Buolamwini and Gebru audited three commercial gender classifiers in 2018: “We evaluate 3 commercial gender classification systems using our dataset and show that darker-skinned females are the most misclassified group (with error rates of up to 34.7%). The maximum error rate for lighter-skinned males is 0.8%.” NIST went wider in 2019 — 189 algorithms from 99 developers, 18.27 million images of 8.49 million people — and found one-to-one false-positive rates higher for Asian and African American faces than for Caucasian faces “from a factor of 10 to 100 times”. No global accuracy figure in either study was doing anything but hiding that.
4. Include uncertainty in the rule itself. The IDx-DR endpoints were tested one-sided at a 2.5% Type I error. So the one-sided 97.5% bound had to clear the threshold, not the point estimate. In 900 participants at 10 primary-care sites the observed values were 87.4% sensitivity and 89.5% specificity, and the abstract states the verdict in the form the protocol demanded: “The AI system exceeded all pre-specified superiority endpoints at sensitivity of 87.2% (95% CI, 81.8–91.2%) (>85%), specificity of 90.7% (95% CI, 88.3–92.7%) (>82.5%), and imageability rate of 96.1% (95% CI, 94.6–97.3%)” — Abràmoff and colleagues, npj Digital Medicine, 2018. Sensitivity came in at 87.2% against a bar of 85%, and the lower bound of its interval, 81.8%, sits below that bar. The specificity margin and the imageability margin are what carried the trial. A gate that read only the three point estimates could not have told you that.
5. Define dispositions. The FDA granted the De Novo request on 11 April 2018 — one disposition among release, limited rollout, revise, collect more data, and reject. A gate that can only say yes is not a gate. The section that follows shows what the missing dispositions cost.
1. State the primary claim
Choose one central outcome or a clearly ordered set.
2. Set minimum guardrails
Define unacceptable harm, workload, fairness, latency, or cost.
3. Add slice floors
Require adequate evidence and performance for critical populations.
4. Include uncertainty
Use confidence bounds or equivalence margins where appropriate.
5. Define dispositions
Specify release, limited rollout, revise, collect data, or reject.
Analogy
A launch constraint that was waived for every flight
The obvious analogy for a release gate is a launch checklist: one failed pressure seal stops the count, whatever the rest of the board says. The instructive part is not that spacecraft have tolerances. It is what happened to a real one.
A primary O-ring failed on STS 51-B. Afterwards a launch constraint was placed on the Shuttle — formally, a condition that had to be resolved before each flight — by Solid Rocket Booster Project Manager Lawrence Mulloy and the Marshall Problem Assessment Committee. The Presidential Commission on the Space Shuttle Challenger Accident recorded in June 1986 what the constraint then did: “After the launch constraint was imposed, Project Manager Mulloy waived it for each Shuttle flight after July 10, 1985.” The Commission also found that NASA Levels I and II did not know the constraint existed.
A second investigating body reached the same conclusion about the mechanism. The House Committee on Science and Technology, in October 1986: “Launch constraints were often waived after developing a rationale for accepting the problem rather than correcting the problem; moreover, this rationale was not always based on sound engineering or scientific principles.”
That is the failure mode a release gate inherits. A pressure seal has a tolerance printed on it. A metric threshold was argued into place rather than measured, which makes it cheap to renegotiate under deadline. And the renegotiation always arrives with a rationale attached, because a waiver that came without one would be refused.
A constraint waived for every flight after July 10, 1985 was documentation, not a veto.
Key idea
Point estimates should not sit exactly on the gate
If a candidate clears its threshold by less than its measurement uncertainty, the honest verdict may be inconclusive. Two drug regulators have written that discipline down. Both forbid the move that release reviews make routinely: setting the margin once the result is known.
Pre-specification has been the European rule since 2005. The European Medicines Agency's guideline on the choice of the non-inferiority margin states that “the recommended approach is to pre-specify a margin of non-inferiority in the protocol”, and it refuses to let that margin be chosen for convenience: “The choice of delta must always be justified on both clinical and statistical grounds.” It warns specifically against defining “an arbitrary achievable delta”. And it reads the confidence interval rather than the estimate: the two-sided 95% interval, equivalently the one-sided 97.5% bound, must lie entirely on the positive side of the margin.
The FDA is blunter. Its 2016 guidance on non-inferiority trials says “the NI margin must be prespecified”, then names the retrofit and rejects it: “an unplanned determination of non-inferiority following failure to show superiority, when the margin was not determined until results of the trial were known, would not be sufficient for demonstrating non-inferiority of the test drug.” Read that as a release rule and it covers the common case exactly. The candidate misses the primary bar, and a threshold for “no worse than the incumbent” is drafted that afternoon.
A marginal pass can still justify a limited rollout with monitoring. That is a different disposition from claiming the bar was confidently cleared. The portfolio should have named both in advance.
A margin chosen after the result is not a margin, and two regulators say so in writing.
Cost-quality frontiers expose dominated models
Plot quality against latency, reviewer hours, memory, or monetary cost. A model is dominated when another option is at least as good on every relevant axis and better on at least one. Dominated candidates should leave the room before anyone debates value tradeoffs.
The axes are rarely hypothetical. For the Epic Sepsis Model both were already measured at Michigan Medicine. Discrimination of 0.63 (95% CI 0.62–0.64) on one axis. On the other, an alert for 6971 of 38 455 hospitalized patients — a number needed to evaluate of 8, which is reviewer hours in the plainest possible units. When a candidate sits low on quality and high on operational cost at the same time, the frontier does the arguing and the stakeholders are spared it.
The same view reveals the opposite case, where a tiny metric gain demands a large operational sacrifice. It makes that sacrifice visible as a quantity rather than as a preference.
Do not negotiate over a model that is worse and more expensive.
Case
One hundred milliseconds, priced at Bing
One hundred milliseconds of delay at Bing is worth 0.6% of revenue. Kohavi and colleagues measured it in 2013, by slowing users down on purpose: “We recently ran a slowdown experiment where we slowed 10% of users by 100msec (milliseconds) and another 10% by 250msec for two weeks. The results showed that performance absolutely matters a lot today: every 100msec improves revenue by 0.6%.”
They then convert that rate into a staffing decision. An engineer who improves server performance by 10msec, they write, “more than pays for his fully-loaded annual costs”. Ten milliseconds — a quantity no quality metric on the dashboard can see — is worth more than a salary. Latency is not a footnote to the metric portfolio. On this evidence it belongs in the same column as the primary outcome, with its own threshold and its own veto.
Figure
Write the decision table before the result table
A decision table names each criterion, unit, population, threshold, uncertainty rule, owner, and resulting action. Written before the results, it makes exceptions visible and stops a dashboard from becoming an invitation to rationalize the preferred model. Metrics remain essential. Governance decides how they become a decision.
What it costs when the only available disposition is ship-to-everyone has a date and a size. On 19 July 2024 a CrowdStrike Rapid Response Content update — a new version of Channel File 291 — went to Falcon sensors with no staged rollout. The Content Validator assessed the new Template Instances expecting the template to be given 21 inputs. The sensor's integration code supplied 20. The resulting out-of-bounds read crashed machines worldwide. The blast radius was measured from outside the company: “We currently estimate that CrowdStrike's update affected 8.5 million Windows devices, or less than one percent of all Windows machines.” — David Weston, Vice President, Enterprise and OS Security at Microsoft.
CrowdStrike's own root cause analysis, dated 6 August 2024, lists the missing gate as Finding 6, “Each Template Instance should be deployed in a staged rollout”, and describes the disposition the release process did not have: “New Template Instances that have passed canary testing are to be successively promoted to wider deployment rings or rolled back if problems are detected.” Canary, ring, rollback — three dispositions between release and reject, absent from the table on 19 July.
The base rate that tells you what a gate is for comes from the same Bing paper. Its third tenet is “We are poor at assessing the value of ideas”, and the evidence is a count: “Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve”. The authors add that “Success is even harder to find in well-optimized domains like Bing”. A release gate that never blocks anything is governing a process in which most proposals should not ship. A gate that refuses is doing the job it was written for.
A portfolio is complete only when it specifies what the organization will do.
Key takeaways
- A metric portfolio assigns distinct roles — primary outcome, guardrail, diagnostic, operational. The Epic Sepsis Model shows what happens without them: AUC 0.63 (95% CI 0.62–0.64) against a claimed 0.76–0.83, 67% of sepsis cases missed, and alerts on 6971 of 38 455 hospitalized patients.
- Guardrails and slice floors must not be averaged away by a strong global score: Buolamwini and Gebru measured error rates up to 34.7% for darker-skinned females against a 0.8% maximum for lighter-skinned males, and NIST found one-to-one false-positive rates differing "from a factor of 10 to 100 times" across 189 algorithms.
- Precommitted acceptance criteria reduce cherry-picking and make exceptions auditable. Thresholds of 85.0% sensitivity and 82.5% specificity were in the IDx-DR protocol before enrolment, and the EMA and FDA both require a non-inferiority margin to be prespecified.
- Pareto analysis exposes candidates that are both worse and more costly than an alternative. It preserves the tradeoff instead of collapsing it into a scalar invented after the winner is known — 29 teams on one dataset produced effects from 0.89 to 2.93 in OR units.
- Uncertainty belongs inside the rule, not the appendix: IDx-DR was judged on a one-sided 97.5% bound rather than its point estimates, which is why its 87.2% sensitivity with a lower bound of 81.8% is a different finding from its 90.7% specificity.
- A release gate is complete only when each measurement maps to a concrete action — canary, ring, rollback, reject. CrowdStrike's Finding 6 of 6 August 2024 reached that conclusion after an update with no staged rollout affected an estimated 8.5 million Windows devices.