Evaluation
Fairness Metrics, Tradeoffs, and Decision Context
Evaluate group fairness through outcome, error-rate, calibration, and exposure measures while confronting base rates, label quality, uncertainty, and incompatible criteria.
By the end you can
- Distinguish demographic parity, equal opportunity, equalized odds, and predictive parity
- Explain why fairness criteria can conflict when base rates differ
- Connect statistical disparities to labels, decisions, processes, and remedies
- Report uncertainty, intersections, and affected-stakeholder perspectives
A parity gap is a finding, not a complete diagnosis
In the 2023 HMDA data, denial rates on conventional first-lien home-purchase loans were 16.6% for Black applicants, 12.0% for Hispanic-White applicants, 9.0% for Asian applicants and 5.8% for non-Hispanic-White applicants. The file covers 5,113 reporting institutions. That is a finding.
The Consumer Financial Protection Bureau publishes those rates. In the same summary it says what they are not: “HMDA data are generally not used alone to determine whether a lender is complying with fair lending laws. The data do not include some legitimate credit risk considerations for loan approval and loan pricing decisions.”
Some of those considerations have been put in. Confidential 2018-2020 HMDA records carry credit score, loan-to-value and debt-to-income, the fields the public file lacks. Ky and Lim re-estimated the same kind of gap on those records at the Federal Reserve Bank of Minneapolis, in a 2022 working paper: “In our baseline specification, we estimate that Black applicants are 2.9 percentage points more likely to have their mortgage application denied relative to White applicants, while Asian and Latinx applicants are 2.2 percentage points and 1.5 percentage points, respectively, more likely to be denied.”
Both halves of that result matter. Underwriting controls narrowed the gap. They did not close it.
A fairness metric describes one statistical relationship. Historical access, the composition of the application pool, label bias, measurement quality, threshold policy and the recorded target all sit upstream of the number. The surrounding decision process is what any remedy has to act on.
Controls narrowed the Black-White denial gap to 2.9 percentage points; they did not close it.
Case
When prevalence differs, a satisfied criterion still produces disparate impact
A model can meet a fairness criterion and still hurt one group. Chouldechova showed the mechanism in 2016, studying recidivism prediction instruments: “adherence to the criterion may lead to considerable disparate impact when recidivism prevalence differs across groups”. The Broward County COMPAS records are what that sentence looks like with numbers in it.
ProPublica built contingency tables for more than 10,000 Broward County defendants. The false-positive rate was 44.85% for Black defendants against 23.45% for white defendants. The false-negative rates ran the other way: 27.99% against 47.72%.
The same data was reanalysed in Federal Probation, published by the Administrative Office of the United States Courts. Flores and colleagues found the instrument about equally accurate for both groups: “The AUC estimate for White defendants was .69 and .70 for Black defendants, with no significant difference between values by race.” Their two-year rearrest base rates were 47% overall, 39% white and 52% Black.
Equal accuracy by AUC-ROC, unequal prevalence, and the two false-positive rates 44.85% and 23.45% sitting on top of each other. Prevalence drives the result and the criterion does the rest. Nobody has to be careless.
Comparison
Common group criteria condition on different events
Each definition expresses a different priority. One of them has been U.S. regulation since 1978.
Demographic parity is the four-fifths rule, codified at 29 CFR 1607.4(D) in the EEOC's Uniform Guidelines on Employee Selection Procedures: “A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact…”
The agencies that wrote the rule also wrote down what it does not do. From their 1979 interpretive Q&A: “The 4/5ths rule of thumb speaks only to the question of adverse impact, and is not intended to resolve the ultimate question of unlawful discrimination.” NIST Special Publication 1270 quotes the same provision when surveying how U.S. courts and regulators test disparate impact.
So the first column below is not an academic construct. It is a threshold of 0.8 with a regulator's own disclaimer attached, and the other three carry the same status. Each conditions on a different event. None of them settles the question by itself.
Demographic parity
Positive decision rates are equal across groups.
- Conditions on group only
- Ignores reference outcome
- Relevant to access or exposure
- Can conflict with utility or other criteria
Equal opportunity
True-positive rates are equal across groups.
- Conditions on positive outcome
- Focuses on missed opportunity
- Needs trustworthy labels
- Does not equalize false positives
Equalized odds
Both true-positive and false-positive rates are equal across groups.
- Conditions on reference label
- Balances error rates
- Can require group-aware thresholds
- May conflict with calibration
Predictive parity
Precision is equal across groups.
- Conditions on positive prediction
- Relates to decision reliability
- Depends on prevalence
- Can conflict with error-rate parity
Visual
Where disparities enter a system
Fairness evaluation should cover more than model outputs. The best-documented demonstration of that entered at step 2. A widely used commercial care-management algorithm predicted health-care cost rather than illness. Cost stood in for need, and the difference between the two rode into the score.
Obermeyer and colleagues dissected the system in Science in 2019: “At a given risk score, Black patients are considerably sicker than White patients, as evidenced by signs of uncontrolled illnesses.” Referral for extra help ran at 17.7% for Black patients where a direct health measure would have referred 46.5%.
The defect was the label, so the fix was the label. Not the threshold, not the model class, not the feature set.
The consequence stage arrived the same day. New York's Department of Financial Services and Department of Health wrote to UnitedHealth Group about the system, Optum's Impact Pro. The Superintendent of Financial Services and the Commissioner of Health told the company's CEO: “These discriminatory results, whether intentional or not, are unacceptable and are unlawful in New York.”
1. Eligibility
Who can enter the process and who is missing from the data?
2. Measurement
How are features and outcomes observed, delayed, or contested?
3. Modeling
Which patterns, objectives, and constraints shape scores?
4. Decision policy
How do thresholds, overrides, capacity, and appeals convert scores into actions?
5. Consequences
Who receives benefit, burden, delay, scrutiny, or exclusion over time?
Example
Equal rates can conceal unequal experiences
A hiring screen is adjusted to equalize true-positive rates across two groups. Whether the people it rejects can ever inspect that adjustment is a separate question. It has been litigated.
Mobley v. Workday is a nationwide ADEA collective action in the Northern District of California. The court describes it as challenging “an AI screening system that is more likely to deny applicants who are African American, suffer from disabilities, or are over forty years old”. The EEOC filed an amicus brief in April 2024 under Title VII, the ADEA and the ADA.
In May 2026 the plaintiffs moved to see the vendor's own fairness work. Magistrate Judge Laurel Beeler ruled against them: “The bias-testing data is privileged.” The order concludes, “The court denies the motion to compel production of Workday’s bias-testing data and its customers’ applicant data.”
Bias testing routed through counsel can be withheld from the applicants the screen turned down.
- Label concern: historical “successful employee” labels reflect unequal mentoring and promotion opportunities, so equalizing true-positive rates equalizes performance against a contested target.
- False positives: equal opportunity does not require equal false-positive rates, so unnecessary interviews may still differ across groups.
- Access: people discouraged from applying never enter the evaluation denominator.
- Intersection: aggregate gender parity can hide failure for older women in one job family.
- Remedy and disclosure: the fix may belong upstream, in how the interview itself is run, rather than in a threshold applied afterwards — and, as Mobley shows, whatever bias testing exists may never reach the people with standing to question it.
Fairness tables need support and intervals
Subgroup rates based on few positive outcomes can swing sharply. Report counts, prevalence, missingness, label delay, confidence intervals, and the randomization or sampling unit.
Do not interpret “not statistically significant” as evidence of equality. Equivalence requires a justified margin and enough precision to rule out disparities that matter.
A wide interval is a data limitation, not a fairness certificate.
Analogy
Balancing several scales with linked weights
Because several scales are strung together, adjusting one balance shifts the others. A designer must decide which balance represents the relevant safety or justice goal.
A scale is trued by moving a weight. A parity gap is not. The string it hangs from runs back through who applied, what was recorded, and which outcome was ever written down. The coupling between criteria is arithmetic. The choice of which one to satisfy is not.
Improving one parity measure can worsen another without any arithmetic mistake.
Key idea
Some desirable criteria cannot all hold simultaneously
Groups have different base rates, and predictions are imperfect. Under those conditions, calibration within groups and equalized error rates generally cannot all be satisfied at once, except in special cases. That is not only a result in the literature. It is instruction from a national standards body.
NIST Special Publication 1270, published in 2022, tells practitioners: “When deciding which fairness metric to adopt, it is important to recognize the impossibility of satisfying certain mathematical fairness constraints at once except in highly constrained special cases.” The reference it cites for that passage is the theorem in the next section. Its own section closes: “The plethora of fairness metric definitions illustrates that fairness cannot be reduced to a concise mathematical definition.”
This is not permission to abandon fairness. It means whoever owns the decision must say which harms, rights, and principles come first, rather than presenting one metric as universally correct.
Metric incompatibility turns fairness into an explicit governance choice.
Case
The theorem that rules out satisfying three fairness criteria at once
Three fairness conditions cannot all hold at once. That is a theorem, not an engineering complaint. Kleinberg, Mullainathan and Raghavan “formalize three fairness conditions that lie at the heart of these debates” and ask whether all three can hold together. Their answer: “except in highly constrained special cases, there is no method that can satisfy these three conditions simultaneously”.
Approximation is not an escape route either — “even satisfying all three conditions approximately requires that the data lie in an approximate version of one of the constrained special cases identified by our theorem”.
The paper appeared in 2016. Better data collection and more careful engineering do not lift the result. That is why NIST passes it to practitioners as guidance rather than as literature.
Steps
Build a fairness evaluation with stakeholders
Statistical analysis should follow a documented harm model. Step 5, evaluating remedies, has a field test with numbers attached. One jurisdiction mandated a parity audit, and then someone counted the results.
New York City's Local Law 144 of 2021 is the first law anywhere to require annual independent bias audits of automated employment decision tools, posted in public. The city's Department of Consumer and Worker Protection states that “DCWP will begin enforcement of this law and rule on July 5, 2023.”
A study presented at FAccT in 2024 counted what was actually posted. Its 155 student investigators checked 391 employers. They found 18 posted audit reports and 13 posted transparency notices. Of what was posted, they write: “Employer discretion may also explain our finding that nearly all audits reported an impact factor over 0.8, a rule of thumb often used in employment discrimination cases.”
Enforcement was counted too. The Office of the New York State Comptroller audited DCWP for July 2023 to June 2025 and reported in December 2025. The agency had received two AEDT complaints in two years and had identified one non-compliance issue. The Comptroller's own review of the same 32 companies found at least 17.
Compare remedies across data, workflow, policy, threshold and model changes using outcome evidence. Hold a published ratio to the same standard as any other intervention: what did it change?
1. Map decisions and harms
Identify benefits, burdens, rights, and recovery pathways for affected groups.
2. Audit data and labels
Examine access, representation, measurement error, and historical processes.
3. Select criteria
Choose metrics that correspond to the stated concern and explain tradeoffs.
4. Analyze intersections
Report critical combinations with support and uncertainty.
5. Evaluate remedies
Compare data, workflow, policy, threshold, and model changes using outcome evidence.
Key takeaways
- A parity gap is a finding, not a diagnosis. In the 2023 HMDA file, denial rates were 16.6% for Black applicants and 5.8% for non-Hispanic-White applicants; credit score, loan-to-value and debt-to-income controls cut the gap to 2.9 points, and it stopped there.
- Demographic parity, equal opportunity, equalized odds and predictive parity condition on different events. The first is codified at 29 CFR 1607.4(D) as an 80% screening threshold, and its own authors say it does not resolve the question of unlawful discrimination.
- Differing base rates make desirable criteria incompatible in practice: COMPAS scored .70 and .69 AUC-ROC across race on two-year rearrest base rates of 52% and 39%, and produced false-positive rates of 44.85% against 23.45%.
- The incompatibility is a theorem from 2016, and NIST Special Publication 1270 hands it to practitioners as guidance — so selecting a metric is a governance act with a citation behind it.
- Disparities enter through labels and access, not only through models: a care-management algorithm that predicted cost instead of illness referred Black patients at 17.7% where a direct health measure would have referred 46.5%.
- Remedies must be judged by outcomes. Local Law 144 yielded 18 posted audit reports across 391 employers with nearly all impact ratios above 0.8, and Mobley v. Workday shows a vendor's bias-testing data can be withheld as privileged.