Evaluation
Slice Analysis and Subgroup Performance
Design slice evaluations that find concentrated failure without creating a garden of noisy, post-hoc comparisons.
By the end you can
- Define slices from intended use, risk analysis, and observed failure mechanisms
- Report performance, support, prevalence, and uncertainty for each slice
- Distinguish predeclared critical slices from exploratory discovery
- Handle intersections and multiple comparisons responsibly
0.35 and 0.19, inside one word error rate
Five commercial speech recognizers — Amazon, Apple, Google, IBM and Microsoft — were given the same 19.8 hours of sociolinguistic interviews. The speakers were matched on age and gender: 42 white, 73 black. The average word error rate came out at 0.35 for black speakers and 0.19 for white speakers. The gap was there in all five. Microsoft's was the narrowest, 0.27 against 0.15. Apple's was the widest, 0.45 against 0.23.
The result went into PNAS in 2020. Koenecke and colleagues put it in one sentence, in the significance statement: “By analyzing a large corpus of sociolinguistic interviews with white and African American speakers, we demonstrate large racial disparities in the performance of five popular commercial ASR systems.”
Any single number reported over that corpus is a blend of 0.35 and 0.19, weighted by how many speakers of each kind happened to be recorded. It describes no speaker in the study. Slice analysis asks where the aggregate came from. It asks whether an important population or condition is being averaged away.
A global metric is a weighted mixture, not a universal experience.
Case
34.7% and 0.8%, inside one reported accuracy
Three commercial gender classifiers were tested on a new set balanced by gender and skin type. Buolamwini and Gebru built that set after finding the two standard benchmarks were 79.6% and 86.2% lighter-skinned. Darker-skinned women were the most misclassified group, at error rates “of up to 34.7%”. For lighter-skinned men the maximum was 0.8%. One average covers both.
Figure
Visual
Where useful slices come from
Good slices are tied to hypotheses about risk or mechanisms. The sharpest ones are usually defined by what the system does to people, not by whatever attribute is convenient to group on.
One of the sharpest was drawn by the product itself. A commercial risk-prediction algorithm auto-enrols patients into a care programme at the 97th percentile of risk score. That threshold is the slice. At it, Black patients had 26.3% more chronic illnesses than White patients: 4.8 distinct conditions against 3.8, P<0.001. The algorithm was predicting cost rather than illness.
The study covered 6,079 self-identified Black and 43,539 self-identified White patients — 11,929 and 88,080 patient-years. Science published it in 2019. Obermeyer and colleagues put the size of it in the abstract: “We show that a widely used algorithm, typical of this industry-wide approach and affecting millions of patients, exhibits significant racial bias: At a given risk score, Black patients are considerably sicker than White patients, as evidenced by signs of uncontrolled illnesses. Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%.”
That single slice is three of the rows below at once: an intended-use population (who is served by the care programme), an operational pathway (auto-enrolment), and severity (uncontrolled chronic illness).
It also shows how fast consequence turns into obligation. The paper appeared on 25 October 2019. The same day, New York's Superintendent of Financial Services, Linda A. Lacewell, and Health Commissioner Howard A. Zucker wrote to UnitedHealth Group. They required it to “immediately investigate these reports and demonstrate that this algorithm is not racially discriminatory or to cease using Impact Pro (or any other data analytics program) if you cannot demonstrate that it does not rely on racial biases or perpetuate racially disparate impacts”. Consequence, not convenience, is what made that the right slice to cut.
Intended-use populations
Regions, languages, age ranges, devices, or customer tiers explicitly served.
Failure mechanisms
Low light, long documents, rare classes, missing features, or domain shift.
Operational pathways
Manual review, automatic action, abstention, or fallback route.
Severity and consequence
High-cost errors, vulnerable users, safety events, or legal obligations.
Exploratory discovery
Data-driven clusters or residual patterns that generate new hypotheses.
Example
Every slice needs more than one score
A subgroup table should carry the context a reader needs to read it. The clearest demonstration of why is a deployment where nobody built the table at all.
Rite Aid ran facial recognition in hundreds of pharmacies with none of these columns in place, the Federal Trade Commission alleges. Exposure was uneven. Roughly 80 percent of Rite Aid stores sit in plurality-White areas. About 60 percent of the stores using the technology were in plurality non-White areas. The primary metric by slice was already sitting in the logs: match alerts in plurality-Black or plurality-Asian areas were significantly more likely to carry low confidence scores. Employees recorded thousands of false-positive match alerts between December 2019 and July 2020, over 5,000 of them in stores more than 100 miles from the store that created the enrollment.
The complaint, filed on 19 December 2023, states the omission plainly: “However, Rite Aid made no effort, either before implementing facial recognition technology or at any time while using the technology, to assess, test, inquire, or monitor whether the accuracy of its facial recognition technology varied depending on characteristics of the image subject, including whether the technology was especially likely to generate false positives depending on image subjects' race or gender.”
The order came on 23 February 2024. District Judge Kelley Brisbon Hodge signed a stipulated order for permanent injunction, barring Rite Aid from facial recognition surveillance for five years. The first party to assemble the subgroup table was the regulator.
- Support: Number of independent units and positive outcomes in the slice.
- Prevalence: Base rate or target distribution, which can change predictive values.
- Primary metric: The quantity linked to intended use for that slice — for Rite Aid's stores, the confidence score carried by each match alert.
- Uncertainty: Confidence interval or posterior summary using the correct analysis unit.
- Exposure: How often the group encounters the system — about 60 percent of the deploying stores stood in plurality non-White areas, against roughly 80 percent of all stores in plurality-White ones.
- Comparator: Incumbent or baseline performance within the same slice.
Comparison
Critical and exploratory slices have different evidentiary roles
Both are useful when reported honestly. One regulatory regime has enforced exactly that division since 5 February 1998, the day ICH guideline E9, on statistical principles for clinical trials, reached Step 4. Its section 5.7 requires planned subgroup analyses to be set out in the protocol in advance. Everything else has to be labelled for what it is: “In most cases, however, subgroup or interaction analyses are exploratory and should be clearly identified as such; they should explore the uniformity of any treatment effects found overall.”
The same section attaches a consequence: “When exploratory, these analyses should be interpreted cautiously; any conclusion of treatment efficacy (or lack thereof) or safety based solely on exploratory subgroup analyses are unlikely to be accepted”. The FDA adopted the guideline in September 1998. There the passage reads: “Any conclusion of treatment efficacy (or lack thereof) or safety based solely on exploratory subgroup analyses is unlikely to be accepted.”
The protocol is the precommitment device this section is asking for. A slice named before the numbers exist can carry a claim. The identical slice found afterwards can only raise one.
Predeclared slices
Chosen before results from intended use and risk analysis.
- Support release gates
- Reduce cherry-picking
- Need adequate sample plans
- Should include intersections
Exploratory slices
Discovered after inspecting errors or residuals.
- Generate hypotheses
- Can find unknown failures
- Need confirmation on fresh data
- Face multiple-comparison risk
Intersections reveal mechanisms and multiply uncertainty
Performance can be acceptable by language and by device separately yet fail for one language on one device. NIST measured that shape at scale. Its December 2019 demographic study of face recognition ran 18.27 million images of 8.49 million people through 189 mostly commercial algorithms from 99 developers. The executive summary reports two very different sizes of effect: “Across demographics, false positives rates often vary by factors of 10 to beyond 100 times. False negatives tend to be more algorithm-specific, and vary often by factors below 3.”
The interactions, not the headline factors, are what a one-dimensional table would have missed. False positives were highest in American Indians on the domestic law-enforcement images, and within that finding “the relative ordering depends on sex”. False negatives were higher for people born in Africa and the Caribbean, but only on the lower-quality border-crossing images, “the effect being stronger in older individuals”. The direction of the effect flipped with the dataset. Demographic attributes interacted with sex, age and image quality rather than acting alone.
The number of possible intersections grows quickly. Prioritize risk-based combinations, use hierarchical models cautiously, and confirm surprising discoveries on fresh evidence.
Intersectional analysis should be guided by plausible mechanisms and consequences.
Key idea
Small slices need humility, not disappearance
A wide interval is evidence of insufficient precision, not evidence that the group does not matter. Pooling can improve stability, but it may hide real differences if the assumptions are wrong. Report counts, uncertainty, and data gaps.
Pulse oximeters show what happens when the minimum is written down and set at two people. FDA's March 2013 510(k) guidance made the primary metric a single pooled figure — ARMS across all measurements from all subjects — with an acceptance target below 3.0% for transmittance sensors. Underneath it, very little slice structure was asked for: “The FDA pulse oximeter guidance recommends that desaturation studies include 10 or more healthy subjects that vary in age and gender, include 200 or more data points (i.e., paired observations of SpO2-SaO2), and for the study subjects to have a range of skin pigmentation, including at least 2 darkly pigmented subjects or 15% of the study group, whichever is larger.” A pooled error statistic clears its threshold. A two-person slice cannot fail one.
Moving that rule took a decade and a journal paper. In 2020 Sjoding and colleagues reported racial bias in pulse oximetry. FDA convened its Anesthesiology Devices Advisory Committee on 1 November 2022. The UK review chaired by Professor Dame Margaret Whitehead reported on 11 March 2024 that “Pulse oximeters overestimate true oxygen levels in people with darker skin tones, which is exacerbated in patients with low levels of oxygen saturation.” FDA's draft guidance followed on 6 January 2025, recommending more participants and both the Monk Skin Tone Scale and calculated individual typology angle.
For critical groups, collect targeted evidence rather than declaring “no significant difference.”
Absence of precise evidence is not evidence of equal performance.
Steps
Run slice analysis without manufacturing findings
Combine precommitment with disciplined exploration. Step 1 is the protocol that ICH E9 requires, written while the results are still unknown. Step 2 is the question FDA's 2013 minimum of two darkly pigmented subjects did not ask: how many units does this slice need before its number means anything? Step 4 is where the discipline of the first three steps earns you the right to look freely.
1. Define critical slices
Use intended use, stakeholders, regulation, and known failure mechanisms.
2. Plan support
Estimate the number of units and outcomes needed for useful precision.
3. Report a fixed table
Include global, critical, and operational slices with consistent metrics.
4. Explore residuals
Search for new clusters of failure while controlling multiplicity.
5. Confirm and act
Validate discoveries on new data and map failures to product changes.
Analogy
A soup whose average temperature hides a frozen corner
Measured at one point after stirring, a large pot reads hot. The soup can be hot overall while one unmixed corner remains cold.
Stirring is what fixes a pot. Nothing stirs a population. The cold corner stays cold, and stays hidden inside the mixture, until it is measured on its own — with its own count and its own interval.
A safe average does not guarantee safe components.
Key takeaways
- Global metrics are weighted mixtures that hide localized failure: 0.35 average word error rate for black speakers against 0.19 for white speakers, in the same five commercial ASR systems.
- Useful slices arise from intended use, risk, mechanisms, pathways, severity, and disciplined exploration — the health risk-prediction algorithm gave up its slice at the 97th-percentile score that auto-enrols patients into the care programme.
- Slice reports need support, prevalence, uncertainty, exposure, and matched comparators; Rite Aid ran facial recognition in stores 60 percent of which stood in plurality non-White areas without ever building that table.
- Predeclared critical slices and exploratory slices serve different evidentiary roles — ICH E9 section 5.7 has required the distinction in writing since 5 February 1998.
- Small-slice uncertainty should motivate data collection rather than a claim of equality: a pooled ARMS below 3.0% and a two-subject minimum let a decade of pulse-oximeter disparity pass.
- Intersectional analysis should prioritize plausible high-consequence combinations and confirm discoveries on fresh data — NIST found false positives varying by factors of 10 to beyond 100 while false negatives varied by factors below 3, with orderings that reversed between datasets.