Responsible AI
Group Performance and Intersectional Evaluation
Design group and intersectional evaluation that accounts for denominators, uncertainty, sample size, context, and operational consequences.
By the end you can
- Explain why intersectional evaluation pairs relevant slices with operational metrics, denominators, uncertainty, context, and consequences
- Distinguish Single aggregate metric, Predefined slice analysis, and Exploratory intersection search
- Identify evidence that connects population dimension to operational consequence
- Design a review that moves from choose relevant slices to decide and monitor
The denominator comes before the gap
Disaggregated evaluation asks how error and service outcomes vary across groups, intersections, and operating contexts. A report should pair group averages with denominators, uncertainty, threshold, prevalence, and consequence. Take the denominator first. A nine-point gap across four thousand cases and the same gap across eleven are different findings. Only one of them supports a decision.
United States employment law has carried a numbered answer to that problem since 1978. The four-fifths rule sets a selection-rate ratio of 80% as the general threshold of evidence for adverse impact. It sits in the EEOC's Uniform Guidelines on Employee Selection Procedures. The same paragraph then qualifies that threshold in both directions, on sample size. Greater differences may not constitute adverse impact where they are “based on small numbers and are not statistically significant”. And where the thin numbers point the other way, the regulation does not stop at silence. It sends the team looking: “Where the user's evidence concerning the impact of a selection procedure indicates adverse impact but is based upon numbers which are too small to be reliable, evidence concerning the impact of the procedure over a longer period of time and/or evidence concerning the impact which the selection procedure had when used in the same manner in similar circumstances elsewhere may be considered in determining adverse impact.”
Read that as an instruction rather than a caveat. It does not tell a team to publish the thin cell. It does not tell them to drop it. It tells them to go and find a denominator — over a longer period, or from the same procedure used the same way somewhere else. Intersectional analysis can reveal harms concealed by broad categories. But sparse cells create unstable estimates and privacy risk. So the answer is rarely one table. It is quantitative slices, targeted data collection, qualitative evidence, and hierarchical or pooled methods where they fit.
A gap reported without its denominator can be neither acted on nor dismissed, and teams usually settle that by dismissing it.
Comparison
Single aggregate metric, Predefined slice analysis, or Exploratory intersection search?
One aggregate metric summarizes. Predefined slices test what somebody thought to ask about. Exploratory search finds the pocket nobody predicted. It also creates, in the same motion, the problem that keeps a striking pocket from counting as a finding.
Search a large number of candidate slices and the arithmetic turns against you. The field's own tooling names that problem and pays for it. Slice Finder automates the search, and its authors put the difficulty plainly in 2019: “One problem with performing many statistical tests (due to a large number of candidate slices) is an increased number of false positives. This is what is also known as Multiple Comparisons Problem (MCP)”. So the tool does not report whichever slice looked worst. It controls false positives with alpha-investing, which bounds the marginal false discovery rate E(V)/E(R) at level alpha. A search that skips that step has not found a disparity. It has drawn a hypothesis out of a hat, in a room full of hats.
Single aggregate metric
Summarizes average performance.
- Useful for broad tracking
- Can hide opposing group effects
- Often dominated by common cases
- Insufficient for high-impact release
Predefined slice analysis
Tests known groups and conditions.
- Supports comparability over time
- Requires relevant taxonomy and sample size
- Can miss emergent intersections
- Should be tied to decisions and harms
Exploratory intersection search
Looks for unexpected pockets of failure.
- Can discover hidden combinations
- Raises multiple-testing and privacy issues
- Needs confirmation on independent data
- Useful as hypothesis generation
Visual
A slice result needs all five stated
Population dimension, task condition, metric, uncertainty, consequence: a slice result means very little until all five are stated together. One short paper carries all five in a single sentence. That is why it is worth reading as a template rather than as a medical curiosity.
A pulse oximeter clips to a finger and reports blood oxygen. In 2020 the New England Journal of Medicine published what happens when that reading is checked against the arterial blood gas underneath it. Sjoding and colleagues took the patients whose clip was showing a reassuring number: “In the University of Michigan cohort, among the patients who had an oxygen saturation of 92 to 96% on pulse oximetry, an arterial oxygen saturation of less than 88% was found in 88 of 749 arterial blood gas measurements in Black patients (11.7%; 95% confidence interval [CI], 8.5 to 16.0) and in 99 of 2778 measurements in White patients (3.6%; 95% CI, 2.7 to 4.7).” In a second, multicentre cohort the same comparison ran 160 of 939 measurements against 546 of 8,795 — 17.0% (95% CI 12.2–23.3) against 6.2% (95% CI 5.4–7.1).
Now lay the five elements over it. The population dimension is self-reported race. The task condition is the device reading 92–96%, not the device in general; the finding lives inside that band. The error metric is occult hypoxemia — an error in one specific direction, not an average deviation. The uncertainty is stated as raw counts and 95% confidence intervals. So a reader can see that the Michigan Black-patient estimate rests on 749 measurements and the White-patient estimate on 2,778. The consequence is the one that makes the rest matter: hypoxemia the clinician never sees, on a reading that told them to look elsewhere. Strip any one of the five and the result stops being actionable.
- 1
Population dimension
Protected, demographic, geographic, linguistic, disability, or contextual group.
- 2
Task condition
Lighting, device, language, channel, institution, or rare scenario.
- 3
Error and service metric
False accept, false reject, delay, abandonment, appeal, or workload.
- 4
Statistical uncertainty
Sample count, confidence interval, multiple testing, and stability.
- 5
Operational consequence
Denial, surveillance, extra burden, safety risk, or reduced service quality.
Steps
Declare the slices before the results
Slices chosen after the results arrive prove very little. Choosing them from harm pathways first is what makes the third step worth reporting.
Step one does not have to be invented from scratch. It has had a citable format since 2019. Nine authors proposed that a trained model ship with a short document recording how it was evaluated, and called it a model card. The abstract already specifies what a slice plan should contain: “Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains.”
Two things in that sentence do the work. The intersectional groups are named as examples, not gestured at — age and race, or sex and Fitzpatrick skin type. And relevance is tied to the intended application domain, which is what stops a slice plan from becoming an exhaustive cross-product nobody can power. Steps two through five then have something fixed to report against.
1. Choose relevant slices
Use harm pathways, stakeholder evidence, and domain knowledge.
2. Select operational metrics
Measure the errors and burdens that matter at the chosen threshold.
3. Quantify uncertainty
Report counts, intervals, instability, and multiple comparisons.
4. Investigate mechanisms
Connect disparities to data, measurement, model, interface, or process.
5. Decide and monitor
Collect evidence, redesign, restrict, or compensate with accountable owners.
Analogy
A city average that hides flooded streets
Average rainfall across a city can be reported correctly while one low-lying neighborhood is under water. The aggregate is accurate. And useless for the decision about where to build protection.
Rain falls on a fixed map. The neighborhoods in a model are made by category, threshold and institution interacting. So the flooded area moves when the threshold moves, and the slice plan below has to be chosen with that in mind.
Disaggregation should reveal operational burden, not merely produce more columns.
Key idea
A court found the breach in the measurement never made
Publishing a large table of subgroup metrics is not fairness governance. Reviewers need a decision rule for uncertainty, small samples, conflicting metrics, and what repair follows a disparity. Identity categories are context-dependent and may be sensitive or unavailable. The organization should justify collecting them, protect the data, and avoid treating categories as fixed biological truths.
The harder end of the same point has already been decided in court. South Wales Police deployed live facial recognition, and in 2020 the Court of Appeal held that the force had breached the public sector equality duty in section 149 of the Equality Act 2010. The case is Bridges. The breach was not a disparity the force had measured and tolerated. It was this: “The fact remains, however, that SWP have never sought to satisfy themselves, either directly or by way of independent verification, that the software program in this case does not have an unacceptable bias on grounds of race or sex.”
The force's answer had been that the manufacturer treated the relevant information as commercially confidential. The court held that this was no answer. The duty is the public authority's own, and it is non-delegable. A deploying institution cannot buy its way out of knowing. An unopened vendor claim is not a slice plan.
The breach was found in the verification never sought, not in a number that came out badly.
Case
34.7 per cent against 0.8 per cent, measured at the intersection
The audit that made this concrete measured the intersections directly. Gender Shades tested commercial gender classification products in 2018 and reported error by skin type and gender jointly, rather than one attribute at a time. Buolamwini and Gebru state the result in their abstract: “darker-skinned females are the most misclassified group (with error rates of up to 34.7%). The maximum error rate for lighter-skinned males is 0.8%.”
Neither figure is visible in an aggregate accuracy number. Neither survives a breakdown that reports gender alone. A darker-skin row pools darker-skinned males with darker-skinned females. A gender row pools across skin type. Both marginal tables are compatible with 34.7 per cent and with a far milder worst case. Neither of them can tell you which one you are shipping.
Figure
Example
18.27 million images, and two error types that behave nothing alike
An overall accuracy figure is the wrong summary, and the largest demographic measurement of face recognition to date shows why. In December 2019 NIST ran 18.27 million images of 8.49 million people through 189 algorithms from 99 developers.
The headline is not that the systems are inaccurate. It is that the two error directions do not move together. False positive rates across demographics often vary by factors of 10 to beyond 100 times. False negatives usually vary by factors below 3. On sex, the report is explicit: “We found false positives to be higher in women than men, and this is consistent across algorithms and datasets. This effect is smaller than that due to race.” False positives were elevated in the elderly and in children too.
- Broad grouping: A single accuracy figure averages over sex, age and the rest at once. NIST found false positives higher in women than in men, and elevated in the elderly and in children — all inside that one average.
- Error type: The spread is not one number. False positives often vary across demographics by factors of 10 to beyond 100 times; false negatives usually vary by factors below 3. An accuracy figure that does not say which error it describes has hidden the difference between those two spreads.
- Operating point: The two error types are traded against each other at the threshold. So a single deployed operating point turns NIST's measured spread into different burdens for different groups, depending on which direction the system is tuned to avoid.
- Small cells: NIST reached 8.49 million people across 189 algorithms. A single deployment's own evaluation set will not. Its intersectional cells stay thin and its intervals stay wide, so the sparse-cell rule has to be written before the numbers arrive.
- Consequence: The two errors land differently. Repeated false rejection creates exclusion and forces manual identity checks. A false positive at the same threshold puts the wrong person in front of whatever the system triggers.
Example
A slice plan somebody can be audited against
Declare the slices in advance. Then decide what a cell holding eleven cases is allowed to claim. New York City has written exactly that pair of artefacts into law. Under the 2023 bias-audit rule implementing Local Law 144, an audit of an automated employment decision tool must compute selection rates and impact ratios separately for sex categories, for race/ethnicity categories, and for intersectional categories of sex, ethnicity and race. It must publish the number of applicants in each category. It must state how many assessed individuals fell into an unknown category.
The sparse-cell rule is written down too, with a threshold and a price: “an independent auditor may exclude a category that represents less than 2% of the data being used for the bias audit from the required calculations for impact ratio. Where such a category is excluded, the summary of results must include the independent auditor's justification for the exclusion, as well as the number of applicants and scoring rate or selection rate for the excluded category.” The small cell may leave the impact-ratio calculation. It may not leave the page.
A written plan still has to be checked. That is the last artefact. The New York State Comptroller audited the city's enforcement of Local Law 144 over July 2023–June 2025. The city's Department of Consumer and Worker Protection had reviewed 32 companies and found one non-compliance issue. The auditors found at least 17 instances of potential non-compliance.
- Slice plan: List predeclared groups, intersections, task conditions, and metrics. Local Law 144's audit does it as sex categories, race/ethnicity categories, and intersectional categories of sex, ethnicity and race, with the applicant count published for each.
- Sparse-cell rule: Define in advance when an estimate requires pooling, targeted collection, qualitative review, or no claim. The 2% exclusion is the model: a written threshold, a stated justification, and the excluded category's applicant count and rate still reported.
- Burden metric: Add delay, fallback, abandonment, and appeal effort to error rates, so the report shows operational burden rather than more columns of accuracy.
- Mechanism review: For one disparity, identify at least three plausible causal pathways to test across data, measurement, model, interface, and process. And expect the check that 32 reviews finding one issue against 17 found elsewhere implies: a plan nobody audits is a plan nobody follows.
An average answers a question nobody asked
Averages conceal the groups a system fails. An evaluation that reports only the average has answered a question nobody asked. NIST needed 18.27 million images to establish that false positives and false negatives move by different orders of magnitude across demographics. Sjoding and colleagues needed the arterial blood gas standing behind the oximeter reading to separate 11.7% from 3.6%. Buolamwini and Gebru needed skin type crossed with gender to find 34.7 per cent. In none of these cases did the disparity appear in the number the system reported about itself.
Define when disaggregated and intersectional evaluation requires the team to redesign, restrict, remedy, or retire the system. And set the floor below that. In Bridges the Court of Appeal found the breach of the section 149 duty in the verification never sought, not in a table that came out badly.
Key takeaways
- Aggregate accuracy hides which error you get: NIST found false positive rates across demographics often varying by factors of 10 to beyond 100 times, while false negatives usually vary by factors below 3.
- Intersectional evaluation reveals combinations concealed by broad categories — Gender Shades found error rates of up to 34.7% for darker-skinned females against a maximum of 0.8% for lighter-skinned males.
- Every subgroup metric needs a denominator, operating point, uncertainty, and consequence. At an oximeter reading of 92 to 96%, Sjoding and colleagues reported 88 of 749 measurements (11.7%; 95% CI 8.5 to 16.0) against 99 of 2,778 (3.6%; 95% CI 2.7 to 4.7).
- Sparse cells require caution, targeted evidence, and privacy protection. The four-fifths rule sends teams to a longer period or to the same procedure used elsewhere, and Local Law 144's audit still publishes the count and rate of any category under 2% that it excludes from the impact ratio.
- Exploratory slice discovery should be confirmed on independent evidence, because searching many candidate slices is the Multiple Comparisons Problem — which is why Slice Finder uses alpha-investing to bound the marginal false discovery rate rather than reporting the worst-looking slice.
- Fairness evaluation matters only when disparities trigger investigation and action. In Bridges the Court of Appeal found a breach of the section 149 public sector equality duty because the force had never sought to satisfy itself, directly or by independent verification, that the software was free of unacceptable race or sex bias.