Responsible AI
Labels, Proxies, and Historical Decision Bias
Assess whether labels and proxy targets reflect the desired construct, historical access, prior decisions, or institutional neglect.
By the end you can
- Explain why labels and proxies must be evaluated as products of observation, access, prior decisions, and construct choices rather than treated as ground truth
- Distinguish Outcome label, Human judgment label, and Proxy label
- Identify evidence that connects construct to action alignment
- Design a review that moves from name the construct to align the action
Visual
Where label and proxy governance enters the lifecycle
Construct, operationalization, observation process, prior intervention, action alignment. The gap between what an organization cares about and the column it trains on opens at the second step. A second gap opens at the fourth, and that one has a name.
Picture a courtroom. A judge decides whether to release a defendant on bail. Whether that defendant then fails to return for their court appearance is recorded — but only for the ones released. That is the selective labels problem, named and formalised by Lakkaraju and four co-authors in 2017. The courtroom is in their abstract: "For instance, in the context of judicial bail decisions, we observe the outcome of whether a defendant fails to return for their court appearance only if the human judge decides to release the defendant on bail."
The fourth box exists to flag the consequence. Where a human decision gates whether an outcome is ever observed, comparing human and machine performance on the recorded cases produces erroneous estimates. The authors do not answer that with counterfactual inference. They propose a contraction method, which exploits the fact that human decision-makers disagree with one another. The defendants the judge held never enter the data. They never leave the problem either.
- 1
Construct
The concept the organization actually cares about.
- 2
Operationalization
The observable variable chosen to represent the construct.
- 3
Observation process
Events, access, recording, and timing that produce the label.
- 4
Prior intervention
Human or automated decisions that shape whether the outcome occurs.
- 5
Action alignment
Whether predicting the label supports a legitimate and useful intervention.
A label is a measurement, not ground truth
A label is an operational measurement or decision outcome, not automatically ground truth. Proxy targets can encode access, enforcement, institutional priorities, and prior discrimination while remaining easy to predict. Predictive accuracy cannot validate a target that measures the wrong thing.
The point is not new, and it is not only technical. Discrimination can enter at the first step of data mining, before any model is fitted, in the definition of the target variable and the class labels. Barocas and Selbst made that argument in the California Law Review in 2016. Creditworthiness, "a good employee" — these sound like pre-existing facts about people. They are artifacts of the problem definition. Target-variable definitions, the authors write, "simply inherit the formalizations involved in preexisting assessment mechanisms".
Label review is therefore an inquiry into a decision somebody made. Who defined the construct. How outcomes become observable. Which decisions intervene before observation. Who remains unlabeled. And whether the target supports the intended action.
"Thus, the definition of the target variable and its associated class labels will determine what data mining happens to find." — Barocas and Selbst, California Law Review, 2016
Case
The health algorithm that predicted cost and called it illness
A commercial risk-prediction algorithm decides which patients are flagged for extra help managing their health. Tools of that class reach, by industry estimates, roughly 200 million people in the United States each year. In October 2019 Obermeyer and three co-authors took one of them apart in Science. The editor's summary names the design fault without ceremony: the algorithm "uses health costs as a proxy for health needs".
The figures make the failure legible. Patients are auto-identified for enrolment at the 97th percentile of risk score. At that same score, Black patients had 26.3% more chronic illnesses than White patients: 4.8 versus 3.8 distinct conditions, P<0.001. Same score, more disease. Simulating a predictor with no Black-White gap raised the share of Black patients among those auto-identified from 17.7% to 46.5%.
The authors put the mechanism in the abstract: "Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%. The bias arises because the algorithm predicts health care costs rather than illness, but unequal access to care means that we spend less money caring for Black patients than for White patients."
The label was easy to predict. It measured the wrong thing.
Figure
Comparison
Outcome label, Human judgment label, or Proxy label?
An outcome label records a real event — a hospitalization, a repayment. It can be concrete and reproducible. It can also arrive late, be censored, or still be shaped by access and policy. A human judgment label, such as an interview rating, captures an assessor's classification. It is rich in context, exposed to inconsistent standards and stereotypes, and in need of guidance and agreement analysis. A proxy label uses an observable substitute for a hidden construct — spending standing in for health need. It is often easier to collect. Its validity depends on mechanism and context, and it can encode structural inequality.
Proxy failure is not confined to health care. In 2016 Lum and Isaac ran a reimplementation of the PredPol algorithm over Oakland Police Department records of drug crimes for 2011, and published what came out in Significance. The locations the algorithm flagged were the neighbourhoods already over-represented in those records. In their words, "black people would be targeted by predictive policing at roughly twice the rate of whites". The 2011 National Survey on Drug Use and Health put drug use at roughly equivalent levels across racial classifications.
Recorded arrests were the available column. Crime was the construct. The distance between the two is the whole finding.
Outcome label
Records a real event.
- Can be concrete and reproducible
- May still be affected by access and policy
- Can arrive late or be censored
- Example: hospitalization or repayment
Human judgment label
Captures an assessor’s classification.
- Can include rich context
- May reflect inconsistent standards or stereotypes
- Requires guidance and agreement analysis
- Example: interview rating
Proxy label
Uses an observable substitute for a hidden construct.
- Often easier to collect
- Validity depends on mechanism and context
- Can encode structural inequality
- Example: spending as health need
Key idea
Easy to predict, and still the wrong target
A highly predictable label can be a poor target. Predictability measures how regular the observation process is. It does not establish construct validity, justice, or usefulness for the intended decision. Prior spending was easy to record and easy to predict. That is why it became the label, and why a model trained on it learned what care had cost rather than who was ill.
In one corner of United States law this is not a principle a lesson has to argue for. It has been regulation since 1978. The Uniform Guidelines on Employee Selection Procedures put the standard on the criterion — the thing success is measured against — not on the accuracy of the test. Scores on a selection procedure may not enter into any judgment used as a criterion measure. And "[a]ll criterion measures and the methods for gathering data need to be examined for freedom from factors which would unfairly alter scores of members of any group". The predictor is not what has to clear that bar. The target is.
Many important constructs have no perfect observable label. Responsible design may combine multiple measures, human review, uncertainty, and limits on automation rather than pretend the proxy is truth.
"Whatever criteria are used should represent important or critical work behavior(s) or work outcomes." — Uniform Guidelines on Employee Selection Procedures, 1978
Steps
Turn label and proxy governance into an operating control
Describe the construct without using the field name. A team that struggles with that first step has usually inherited the label and every assumption attached to it. Then trace label production: access, observation, human decisions, timing, missing cases. Then test validity. Then examine how disagreement and proxy failure are distributed across groups and contexts. Then align the action, confirming that predicting the target supports a legitimate intervention.
The third step has machinery behind it. Some constructs cannot be observed at all — socioeconomic status, teacher effectiveness, risk of recidivism. Each has to be operationalised through a measurement model, and every measurement model rests on assumptions. Jacobs and Wallach brought that apparatus across from the quantitative social sciences into fairness work in 2021. What they supply is a way to test the assumptions: construct reliability and construct validity as things a team can check, rather than adjectives a team applies to its own target.
Their abstract says what goes wrong when nobody checks: "This process, which necessarily involves making assumptions, introduces the potential for mismatches between the theoretical understanding of the construct purported to be measured and its operationalization. We argue that many of the harms discussed in the literature on fairness in computational systems are direct results of such mismatches."
Not a bug in the model. A mismatch in the definition.
1. Name the construct
Describe the concept and decision without using the available field name.
2. Trace label production
Map access, observation, human decisions, timing, and missing cases.
3. Test validity
Compare the label with domain evidence and alternative measures.
4. Examine distribution
Measure disagreement and proxy failure across groups and contexts.
5. Align the action
Confirm that predicting the target supports a legitimate intervention.
Example
Where the label came from, and who never appears in it
Label genealogy takes an afternoon. It answers what a model card rarely does: who created this column, under which policy, and for what. Expect to recover a negotiation rather than a decision.
That expectation comes from six months inside a corporate data science team, watching how a problem actually gets specified. Passi and Barocas reported the fieldwork in 2019. The choice of target variables and proxies, they found, is always negotiated and elastic, and rarely worked out with explicit normative considerations in mind. Their abstract puts the consequence to anyone reviewing a model: "Whether we consider a data science project fair often has as much to do with the formulation of the problem as any property of the resulting model."
- Label genealogy: Document who created the label, why, from which process, and under which policy. Expect what Passi and Barocas found in six months inside one team — a negotiated, elastic history rather than a single documented choice.
- Unobserved cases: Identify the people whose outcomes stay unknown because they lacked access or were never selected. This is the selective labels problem, and no performance figure computed on recorded cases can reach them.
- Alternative target: Propose another measure or composite that better reflects the construct. In the health study, a simulated predictor with no Black-White gap moved the share of Black patients auto-identified from 17.7% to 46.5%.
- Intervention check: Ask whether a perfect prediction of the label would improve the desired outcome. United States employment regulation has been asking the hiring version of that question since 1978.
Example
Label and proxy governance under operational pressure
A badly chosen target has already drawn a regulatory response, not only academic criticism. The Science paper appeared on 25 October 2019. Two New York regulators wrote to UnitedHealth Group the same day. Their subject was Optum's data analytics program Impact Pro. What they objected to was the variable it predicts: past health spending. Not the model's code.
- The named parties: Linda A. Lacewell, Superintendent of the New York State Department of Financial Services, and Howard A. Zucker, Commissioner of the New York State Department of Health, writing to David S. Wichmann, CEO of UnitedHealth Group, on 25 October 2019.
- The target under objection: Optum's data analytics program Impact Pro, and specifically its use of past health spending as the variable to predict.
- The mechanism the regulators named: "By relying on historic spending to triage and diagnose current patients, your algorithm appears to inherently prioritize white patients who have had greater access to healthcare than black patients."
- The demand: the letter called on the company to "immediately investigate these reports and demonstrate that this algorithm is not racially discriminatory or to cease using Impact Pro (or any other data analytics program)".
- The legal position: "These discriminatory results, whether intentional or not, are unacceptable and are unlawful in New York."
A perfect prediction of the wrong target
A perfect prediction of the wrong target is still the wrong system. That is why the review ends by asking what the prediction is for. The health algorithm was accurate — accurate about health care costs. That accuracy carried a 26.3% chronic-illness gap at the 97th percentile of risk score into an enrolment decision, for a class of tools reaching roughly 200 million people in the United States each year. Accuracy was the delivery mechanism.
Record which validity failures force the team to redesign, restrict, remedy, or retire the system. The Impact Pro demand of 25 October 2019 came down to two of those four, and no third option: demonstrate the algorithm is not racially discriminatory, or cease using it.
Key takeaways
- A label is an operational outcome or judgment, not automatically objective truth. Barocas and Selbst place discrimination in the definition of the target variable and the class labels — the first step of data mining, before any model is fitted.
- Proxy targets reproduce historical access, enforcement, and institutional priorities. A commercial algorithm that predicted health care costs rather than illness left Black patients with 26.3% more chronic illnesses than White patients at the same 97th-percentile risk score.
- Predictability does not establish construct validity or fairness. Since 1978, United States employment regulation has put the examination on the criterion measures themselves, not on the predictor's accuracy.
- Selective labels arise when prior decisions determine which outcomes become observable. Where a judge decides who is released, performance computed on the recorded cases produces erroneous estimates.
- Target design should be tested against the action the system is meant to support. On 25 October 2019 New York's regulators demanded that UnitedHealth Group either demonstrate Impact Pro was not racially discriminatory or cease using it, on the strength of the target alone.
- Imperfect constructs may require multiple measures, uncertainty, and limits on automation. Jacobs and Wallach make construct reliability and construct validity testable assumptions rather than adjectives.