Skip to content
AI.info

How machines learn

Labels, Targets, and the Limits of Ground Truth

Understand where targets come from, how labels differ from reality, and how disagreement, delay, and policy shape what a model learns.

By the end you can

Key idea

The phrase “ground truth” can hide a chain of decisions

A chargeback is not identical to fraud. A hospital code is not identical to disease, a user click is not identical to satisfaction, and each of the three is an observed signal produced by a process with delays, incentives, missing cases, and definitions.

The hospital code has been audited, and the audit priced the gap. In US Medicare Advantage the label “this beneficiary has condition X” is produced by a process with money attached to it. The HHS Office of Inspector General went looking for where those diagnoses came from. Some appeared only on chart reviews and health risk assessments, on no other service record. In September 2021 the auditors put that group at an estimated $9.2 billion in risk-adjusted payments for 2017. Twenty of 162 MA companies drove a disproportionate share. One company alone accounted for 40% of it, $3.7 billion, while enrolling 22% of MA beneficiaries. The Report in Brief names the mechanism without hedging: “This may create financial incentives for MA companies to make beneficiaries appear as sick as possible.”

MedPAC reached the same place by a different route. Its March 2024 report to Congress estimated that higher MA coding intensity left 2024 risk scores about 20% above those of similar fee-for-service beneficiaries. After the statutory 5.9% coding adjustment they were still about 13% higher. The projected cost was $50 billion in higher payments for 2024.

Calling a label “truth” can be convenient. It should not end the investigation. A model learns the label-generation process as well as the phenomenon the team hopes to predict.

Labels are evidence about reality, not reality without mediation.

Case

Ninety days past due: where a bank's default label was written down

European banking law had to write one of those definitions down in full. Article 178(1)(b) of Regulation (EU) No 575/2013 fixes the moment of default. It has occurred when “the obligor is past due more than 90 days on any material credit obligation to the institution, the parent undertaking or any of its subsidiaries”.

Ninety is not the only possible number, and the law says so. Competent authorities may “replace the 90 days with 180 days for exposures secured by residential or SME commercial real estate in the retail exposure class, as well as exposures to public sector entities”.

How much past due counts is legislated too. A delegated regulation from 2017 sets an absolute component of the threshold that “shall not exceed 100 EUR” for retail exposures and “shall not exceed 500 EUR” for other exposures. Beside it sits a relative component, which “shall be between 0 % and 2,5 % and shall be set at 1 %” of the obligor's total on-balance-sheet exposures unless the competent authority determines otherwise.

No bank found those numbers in its data. Supervisors chose them. A model trained on that column is predicting the choice.

Comparison

Where teaching signals come from

Different label sources create different strengths and failure modes. A direct observed outcome records the quantity the task genuinely concerns, and usually arrives late. A human annotation captures nuance and needs guidance and calibration. A weak or programmatic label scales cheaply and carries the biases of whatever rule or source wrote it. The proxy is the one that looks free, and it is the row with the case law.

A commercial risk-prediction algorithm used to manage the health of populations had to score “health need”. The label it was actually taught was future health-care cost. Obermeyer and colleagues dissected it in Science in 2019. At the 97th percentile of risk score — the point at which patients are auto-identified for the care programme — Black patients had 26.3% more chronic illnesses than White patients: 4.8 against 3.8 distinct conditions, P < 0.001. Predicting active chronic conditions instead of cost would have raised the Black share of those auto-identified from 17.7% to 46.5%. Nothing in the machinery had to change for that. Only the column it was taught to predict. The abstract ends with the general lesson: “We suggest that the choice of convenient, seemingly effective proxies for ground truth can be an important source of algorithmic bias in many contexts.” Regulators moved the same day. On 25 October 2019 the New York State Department of Financial Services and Department of Health wrote jointly to UnitedHealth Group about the same product, Optum's Impact Pro.

Medicine had already run the proxy question as a randomised trial. The Cardiac Arrhythmia Suppression Trial chose encainide and flecainide because they suppressed ventricular premature beats after myocardial infarction, a strong predictor of sudden death. Of 2,309 patients recruited, 1,727 — 75% — had their arrhythmia suppressed and were randomised. Over an average of 10 months the drugs did to the proxy exactly what they had been selected to do. The outcome it stood for moved the other way: “Encainide and flecainide accounted for the excess of deaths from arrhythmia and nonfatal cardiac arrests (33 of 730 patients taking encainide or flecainide [4.5 percent]; 9 of 725 taking placebo [1.2 percent]; relative risk, 3.6; 95 percent confidence interval, 1.7 to 8.5).” Total mortality was 56 of 730 against 22 of 725, 7.7% against 3.0%, relative risk 2.5. Those two arms were discontinued in April 1989.

A proxy that tracks the goal in historical data can still move against it once something starts optimising the proxy. Evidence of proxy validity is not a formality on a checklist. It is the difference between the two studies above and a system nobody audited.

FigureComparison · 4 columns

Direct observed outcome

A later event records the quantity the task genuinely concerns.

  • Often objective but may arrive late
  • Can still be missing or censored
  • Definitions require a time horizon
  • Example: a machine fails within seven days

Proxy outcome

An available signal stands in for a harder-to-measure goal.

  • Faster or cheaper to collect
  • May reward the wrong behavior
  • Needs evidence of proxy validity
  • Example: clicks used as a proxy for usefulness

Human annotation

People judge content, quality, intent, or category.

  • Can capture nuanced criteria
  • Requires guidance and calibration
  • Disagreement may be informative
  • Example: reviewers mark abusive language

Weak or programmatic label

Rules, distant sources, or other models generate noisy supervision.

  • Scales cheaply
  • Carries rule and source biases
  • Quality varies by slice
  • Example: label sentiment from star ratings

Visual

Three concepts that deserve separate names

Clear language helps teams notice when their model is learning a measurement process rather than the intended concept. The phenomenon of interest is the underlying state people care about. The target definition is the operational rule that says what will be predicted and over what horizon — the 90 days, the 100 EUR, the choice of cost rather than chronic conditions. The recorded label is the value that actually lands in the table, after observation, annotation, or programmatic generation. Confusing the three is how a documented policy decision gets discussed as if it were a fact of nature.

FigureLayers · 3 layers
  1. 01

    Phenomenon of interest

    The underlying state or outcome people care about, such as fraud, quality, or need.

  2. 02

    Target definition

    The operational rule that says what should be predicted and over which horizon.

  3. 03

    Recorded label

    The value actually stored after observation, annotation, or programmatic generation.

Disagreement may reveal ambiguity, not laziness

Two qualified annotators can disagree. The instructions may be vague, the example may lack context, or the concept may have a genuine boundary case. Majority vote produces one label but can erase uncertainty that matters later.

How much gets erased has been measured, with annotators nobody would call careless. Breast biopsy slides went to 115 pathologists in active practice across eight US states, who returned 6,900 individual interpretations. Three experienced experts had already agreed a consensus reference diagnosis for each slide. Agreement with it was 75.3% overall.

That single figure hides the shape of the problem. Concordance ran at 96% for invasive carcinoma and 87% for benign without atypia. For atypia it was 48% (95% CI, 44%–52%), with 17% of interpretations overinterpreted and 35% underinterpreted. Disagreement also rose with denser breast tissue, lower pathologist caseload and smaller practice setting. Elmore and colleagues put the conclusion this way: “In this study of pathologists, in which diagnostic interpretation was based on a single breast biopsy slide, overall agreement between the individual pathologists' interpretations and the expert consensus–derived reference diagnoses was 75.3%, with the highest level of concordance for invasive carcinoma and lower levels of concordance for DCIS and atypia.”

A project reporting one agreement number would have reported 75.3%. It would never have seen the category where specialists reading the same slide agreed less than half the time. Measure agreement by category and slice, inspect examples with disagreement, and decide whether the task needs clearer guidance, multiple labels, an “uncertain” option, or a different target. Perfect agreement is not always realistic or desirable.

Case

ChaosNLI: a hundred annotators per item, and no gold answer

Some sentence pairs have no answer a hundred people will agree on. ChaosNLI, collected in 2020, measured how many. It holds 464,500 annotations, a hundred per example. It covers 3,113 items from SNLI and MNLI and 1,532 from αNLI. Every one of those items already shipped with a single “gold” label.

Nie and colleagues, who built it, report that “high human disagreement exists in a noticeable amount of examples in these datasets”. Models, they add, “can barely beat a random guess on the data with low levels of human agreement, which compose most of the common errors”. Optimising accuracy there means optimising toward a label a hundred people could not agree on.

Example

A labeling guide is part of the model specification

A content-moderation project becomes more reliable when the guide records decisions. Those decisions would otherwise live only in reviewers' heads.

At Google the guide was made executable. On three production classification tasks the engineers did not hand-label at all. They wrote labelling functions that reused knowledge already sitting in the organisation: heuristics, taxonomies, existing models. A generative model reconciled the disagreements among them into probabilistic training labels. The system was Snorkel DryBell, published in 2019. The classifiers it produced were “of comparable quality to ones trained with tens of thousands of hand-labeled examples”. Converting non-servable organisational resources into servable models bought “an average 52% performance improvement”.

In that arrangement the labelling policy is version-controlled code. It can be diffed, reviewed, and rerun against last month's data.

  • Scope: distinguish direct threats, quoted threats, satire, news reporting, and fictional dialogue.
  • Context: specify what conversation history or metadata annotators may use.
  • Uncertainty: allow “insufficient context” instead of forcing every case into a confident class — the atypia category in the pathology study is what forcing looks like at 48% concordance.
  • Escalation: define which categories require specialist or legal review.
  • Calibration: run shared batches and discuss disagreements before full production.
  • Versioning: record when policy definitions change so labels from different periods are not treated as identical.

Some outcomes have not had time to happen

A customer observed for five days cannot yet receive a reliable label for “churn within 60 days.” A loan still being repaid has an unresolved long-term outcome. Treating unresolved cases as negatives corrupts the target.

Clinical-trial regulation writes this distinction into binding text, and gives the target definition a name of its own. The name is the estimand. ICH E9(R1), the addendum on estimands adopted in November 2019, defines it as “a precise description of the treatment effect reflecting the clinical question posed by the trial objective”. That is an object separate from any recorded measurement. It is assembled from named attributes, not inherited from whatever the database happened to store.

The guideline then separates two things that look alike in a data table. An outcome can change meaning because something happened: a subject switching treatment is an intercurrent event. Or the outcome can simply not have happened yet. A subject for whom no outcome event can be observed because the trial has ended is the second case: “The latter is administrative censoring which needs to be addressed as a missing data problem in the statistical analysis.”

Possible responses include waiting, choosing an earlier outcome, excluding immature examples, or using methods designed for censored time-to-event data. The right choice depends on the decision. It also depends on the cost of delayed feedback.

Absence of an observed event is not always evidence that the event will never occur.

Analogy

An analogy: several witnesses describing one event

Several witnesses are asked to describe a brief incident, each from a different position. Their reports may overlap, conflict, or omit details. A careful investigator records those differences instead of declaring the first statement absolute truth.

Human and proxy labels work similarly: each is produced from a viewpoint and a process, while some machine-learning targets are directly measured quantities, where witness testimony is nearly always interpretive.

Steps

Build a labeling process that can learn from its own mistakes

Label quality improves through an explicit operating loop, not a one-time instruction document. Define the construct in writing, including the edge cases left unresolved. Pilot a shared sample, so several reviewers label the same varied examples. Then analyse the disagreement, separating instruction problems, missing context, random mistakes, and genuine ambiguity — the 48% atypia figure is what the last category looks like when it is measured rather than averaged away. Revise definitions, examples and escalation rules, then test again. Then monitor production, auditing slices, reviewers, policy versions, and model-selected examples over time.

FigureProcess · 5 steps
  1. 1. Define the construct

    Write what belongs, what does not, and which edge cases remain unresolved.

  2. 2. Pilot a shared sample

    Have multiple reviewers label the same varied examples.

  3. 3. Analyze disagreement

    Separate instruction problems, missing context, random mistakes, and genuine ambiguity.

  4. 4. Revise and calibrate

    Update examples, definitions, and escalation rules, then test again.

  5. 5. Monitor production

    Audit slices, reviewers, policy versions, and model-selected examples over time.

Key idea

A model can change the labels it later receives

After deployment, model outputs may influence which cases are reviewed, investigated, purchased, approved, or recorded. Future labels then reflect the policy created by the model. They are not a neutral sample of reality.

Predictive policing is the worked example. There the recorded label is an arrest, and the model decides where the officers stand. Kristian Lum and William Isaac ran the published PredPol algorithm over Oakland police records of drug crime in 2016. It would send officers back to the same non-white, low-income neighbourhoods. Black people would be targeted at roughly twice the rate of White people, and people of other racial classifications at about 1.5 times. Their estimates from the National Survey on Drug Use and Health put drug use at roughly equivalent rates across racial classifications.

Two years later the mechanism was modelled formally. Ensign and colleagues proved that discovered-crime data fed back into the model produces a runaway loop, and showed that using reported incidents attenuates but cannot remove it. Their abstract states the finding: “Such systems have been empirically shown to be susceptible to runaway feedback loops, where police are repeatedly sent back to the same neighborhoods regardless of the true crime rate.” The loop is a property of the arrangement, not a bug in one implementation.

Teams should mark whether a label arose under model influence, random sampling, a legacy process, or a human override. Otherwise retraining can reinforce blind spots created by earlier predictions.

Once a model affects observation, label data becomes part of a feedback system.

Key takeaways