Skip to content
AI.info

How machines learn

Building a Dataset That Represents the Job

Learn how collection, sampling, coverage, duplication, and documentation determine which population and conditions a model can learn about.

By the end you can

Key idea

79.6% lighter-skinned, and nobody had counted

Two of the standard benchmarks in face analysis were large, public and widely used, and nobody had measured what was in them. Joy Buolamwini and Timnit Gebru went and counted. Their 2018 result: “We find that these datasets are overwhelmingly composed of lighter-skinned subjects (79.6% for IJB-A and 86.2% for Adience) and introduce a new facial analysis dataset which is balanced by gender and skin type.”

Then they used that balanced dataset to score three commercial gender classifiers. Error on darker-skinned women reached 34.7%. On lighter-skinned men the worst of the three reached 0.8%. Same three products, same task. One number in the low tens of a percent, one below one percent. What stands behind both is the composition of the data the systems were built and measured on.

Sample size reduces some forms of uncertainty. It does not repair a collection process that systematically excludes important cases. IJB-A and Adience were not small. Their imbalance was not visible in any accuracy figure published before somebody counted the subjects.

Representativeness is about how examples were selected, not how impressive the row count looks.

Case

250,000 responses a week, and an estimate 17 points wrong

Three surveys of US COVID-19 vaccine uptake were checked against the CDC's benchmark, between 9 January and 19 May 2021. The biggest one was the furthest wrong. Delphi–Facebook ran about 250,000 responses a week. By May it overestimated first-dose uptake by 17 percentage points. The Census Household Pulse survey ran about 75,000 responses every two weeks, and overestimated it by 14. An Axios–Ipsos online panel of about 1,000 responses a week, “following survey research best practices”, tracked the benchmark. Bradley and colleagues published the comparison in Nature on 8 December 2021, and stated the arithmetic plainly. “A survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10.”

This is not one research group's reading of somebody else's survey. The agency that produced the benchmark published the same gap in its own numbers. The CDC's AdultVaxView report on the Household Pulse Survey records: “As of May 4, 2021 (mid-point of the HPS data collection period used in this report), CDC COVID Data Tracker indicated that 56.4% of adults aged ≥18 years had received at least one dose of COVID-19 vaccine, 18.2 percentage points lower than the HPS estimate of 74.6%”. Three quarters of American adults, said the survey. Just over half, said the count of doses actually administered.

The survey was weighted, and the report is explicit that weighting is not a repair: “bias in estimates will remain”. Weights adjust for the characteristics you thought to measure. They cannot reach the thing that made a person answer or not answer in the first place.

Figure

Sample size shrinks only the error that comes from the luck of the draw; against the error introduced by who was able to answer it buys nothing at all.

Visual

Three populations must be named

A dataset is interpretable only when the team distinguishes who or what it wants to serve from what its collection process can actually observe. Target population, sampling frame and observed sample are three different things. Only the third one is in the file.

The gap between them does not stay hidden. It is legible afterwards, in the finished product. NIST's Face Recognition Vendor Test ran 189 mostly commercial algorithms from 99 developers over 18.27 million images of 8.49 million people, and reported in December 2019. False positive rates varied across demographic groups by factors of 10 to beyond 100. The direction of the error tracked where the system had been built: “A number of algorithms developed in China give low false positive rates on East Asian faces, and sometimes these are lower than those with Caucasian faces.”

None of those 189 algorithms shipped with a description of its sampling frame. At 8.49 million people, the frame described itself anyway — in which faces each system confused with which.

FigureLayers · 3 layers
  1. 01

    Target population

    The future entities, events, locations, and conditions for which the system is intended.

  2. 02

    Sampling frame

    The portion of that population reachable through available systems and collection methods.

  3. 03

    Observed sample

    The actual records that remain after participation, logging, filtering, missingness, and labeling.

Comparison

How examples enter the dataset changes what can be learned

Sampling strategies are tools, not universal guarantees. A selection rule can do measurable damage while every person involved is being careful and publishing their method.

In 2019 the CIFAR-10 and ImageNet test sets were rebuilt from scratch, by re-running the original collection processes. Recht and colleagues report what happened next: “We evaluate a broad range of models and find accuracy drops of 3% - 15% on CIFAR-10 and 11% - 14% on ImageNet.” The obvious reading was that years of models had been quietly overfitted to the old test sets.

Then the replication itself was examined. Engstrom and colleagues showed in 2020 that the re-collection's own selection rule had introduced statistical bias. Once that was corrected, about 3.6% of the original 11.7% ImageNet drop remained unexplained. Roughly two thirds of a headline result about models turned out to be a property of how the second dataset had been assembled. The mechanism that decides which examples enter is not a neutral pipe. It is part of the measurement.

FigureComparison · 4 columns

Random sample

Units are selected using a known random mechanism from a defined frame.

  • Supports population estimates under assumptions
  • Requires a usable sampling frame
  • Can still miss rare cases by chance
  • Example: random orders from all regions

Convenience sample

Examples are collected because they are easy to access.

  • Fast and inexpensive
  • Often reflects participation or platform bias
  • Generalization claims must be narrow
  • Example: voluntary app feedback

Stratified sample

Sampling is controlled across known groups or conditions.

  • Improves coverage of important slices
  • Needs reliable stratum definitions
  • May require weighting for population metrics
  • Example: equal review batches by language

Model- or policy-selected sample

Existing rules or scores determine which cases receive labels.

  • Focuses scarce review capacity
  • Can hide unselected failures
  • Creates feedback and verification bias
  • Example: investigators review only high-risk alerts

Steps

Turn “representative” into a table

A coverage matrix makes broad claims testable. Its two hard steps — compare the sample against deployment, then price the empty cells — have both been carried out in public, with the numbers printed.

Step 3, done for real: in 2019 DeVries and colleagues scored six public object-recognition systems on the Dollar Street dataset, 135 classes photographed in 264 homes across 54 countries. “In particular, the accuracy of the systems is approximately 15% (absolute) higher on household items photographed in the United States than it is on household items photographed in Somalia or Burkina Faso.” Accuracy also ran about 10 points lower for households under US$50 a month than for those over US$3,500. Soap is soap. The systems disagreed by 15 points depending on whose bathroom it was in.

Step 4 has a regulator's version, with a minimum written into the cell. The FDA's 2013 guidance on pulse oximeters, issued 4 March 2013, set the calibration sample supporting clearance at 10 or more healthy subjects, and specified: “Your study should have subjects with a range of skin pigmentations, including at least 2 darkly pigmented subjects or 15% of your subject pool, whichever is larger.” Two out of ten counts as a filled cell.

Then the filled cell met patients. Wong and colleagues examined 87,971 patients in 215 hospitals and 382 ICUs, and published in JAMA Network Open on 3 November 2021. They found hidden hypoxemia — arterial oxygen low while the oximeter read acceptable — in 6.9% of Black patients against 4.9% of White patients, associated with mortality and later organ dysfunction. A cell can be occupied and still be empty of evidence.

Step 5 also has a worked instance. The response to the geographic hole was to go and collect the missing rows. A curated, fully labelled Dollar Street of 38,479 images was published in 2022, from homes with incomes as low as $26.99 per month, built precisely because existing datasets under-represent them.

FigureProcess · 5 steps
  1. 1. List important axes

    Choose time, geography, device, language, product, severity, and other task-relevant conditions.

  2. 2. Count entities and examples

    Report both row volume and distinct units for each slice.

  3. 3. Compare deployment

    Estimate how expected live proportions differ from the sample.

  4. 4. Find empty cells

    Identify combinations with little or no evidence.

  5. 5. Choose a response

    Collect more data, narrow scope, adjust sampling, or add a fallback.

Repeated content can create false confidence

Exact duplicates may arise from retries, copied records, mirrored files, or repeated imports. Near-duplicates include frames from the same video, templated messages, slightly cropped images, or multiple windows around one event.

Duplicates can dominate training. They can cross partition boundaries. Deduplication should use domain-aware identity and similarity checks, while preserving legitimate repeated events when recurrence itself matters.

Two of the most-used benchmarks in computer vision were partly testing models on pictures those models had already seen. Barz and Denzler published the count in the Journal of Imaging in 2020. Of CIFAR-10 and CIFAR-100 they write: “We find that 3.3% and 10% of the images from the test sets of these datasets have duplicates in the training set.” They then rebuilt the test sets. The fresh images came from the same domain. The training sets were left untouched, so pre-trained models stayed valid. They re-scored a range of standard architectures. The result was “a significant drop in classification accuracy of between 9% and 14% relative to the original performance on the duplicate-free test set”. Years of published comparisons had been decided partly on images the models had already been trained on.

Missing groups may reflect the system, not the world

People who cannot access a service, devices that fail before logging, and cases resolved outside the platform may never appear. Their absence can be a direct consequence of product design or institutional policy.

Sometimes the rule that removes them is written down, and someone measures what it removed. C4, the 750GB corpus behind T5, was cleaned by discarding “any page that contained any word” on a public list of bad words. Raffel and colleagues set the heuristic out in 2020, in the Journal of Machine Learning Research.

A year later somebody measured whose English that one line deleted. Dodge and colleagues, in 2021: “Using the most likely dialect of a document, we find that AAE and Hispanic-aligned English are removed at substantially higher rates (42% and 32%, respectively) than WAE and other English (6.2% and 7.2%, respectively).” Nearly half of the African American English, against roughly one document in sixteen of the White-aligned English. The filter was not even accurate on its own terms. Of a 100,000-document sample of what it excluded, only 31% was sexual in nature.

Nobody wrote a rule against a dialect. Somebody wrote a word list, and the word list did it. Dataset review should therefore ask who cannot generate a record, who declines participation, which failures stop telemetry, which filters ran before you saw the file, and which outcomes remain unobserved. These questions often require domain and user research. SQL alone is not enough.

Position

Scale answers a question nobody asked

Row count is the one number a data team can produce on demand. So it is the number that reaches the slide. A beginner reasonably concludes that more rows mean more confidence. More rows buy one thing: less error from the luck of the draw. Against error introduced by the selection rule they buy nothing at all. That second error does not shrink as the dataset grows. It is the same at a thousand rows and at a million.

The difference has been priced in about as clean a setting as exists: three surveys, one benchmark, one year. Delphi–Facebook, at roughly 250,000 responses a week, was 17 percentage points high by May 2021. Census Household Pulse, at roughly 75,000 every two weeks, was 14 high. The CDC's own report puts the same gap at 18.2 points on 4 May 2021 — 74.6% claimed, against 56.4% of adults actually recorded as having had a first dose. An Axios–Ipsos panel of about 1,000 a week tracked the benchmark. The arithmetic is the part to keep. “A survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10.”

That ratio is the whole lesson in one line: 250,000 responses, the precision of 10. So “how much data do we have” is a question with no useful answer. The honest move is to stop asking it. Ask instead who could not have entered the dataset at all. Count distinct entities rather than rows. That is why a coverage table, a dataset record and a deliberately narrowed claim keep reappearing in this lesson. Not one of them can be produced by collecting more of what you already have.

Sample size shrinks the error that comes from luck. It does nothing to the error that comes from who was able to answer.

Analogy

An analogy: learning a language from one shelf

Someone tries to understand an entire language by reading only technical manuals donated by one company. The collection may contain millions of sentences. Its vocabulary, tone, authors, and topics stay narrow.

That shelf exists, and it has been measured. Six researchers at Google Brain geo-located roughly 2 million of Open Images' 9 million images, and published the breakdown in 2017. More than 32% were US-based. 60% came from the six most represented countries of North America and Europe. “Meanwhile, China and India – the two most populous countries in the world – were represented with only 1% and 2% of the images, respectively.” Their ImageNet sample was about 45% US-based.

Nine million images is not a shelf anyone would call narrow, and that is the point of the analogy. A dataset can be rich in volume and poor in coverage at the same time. Machine-learning examples are narrowed by more than content variety: labels, temporal drift, measurement systems, and policies all distort them.

Example

Document the dataset before it becomes folklore

A short dataset record can prevent hidden assumptions from spreading between teams.

The practice has a citable origin. Gebru and colleagues proposed “datasheets for datasets” in 2018, and the proposal was published in Communications of the ACM in 2021. The analogy is exact and deliberate. In the electronics industry every component ships with a datasheet describing “its operating characteristics, test results, recommended uses, and other information”. They propose that every dataset be accompanied by one recording “its motivation, composition, collection process, recommended uses, and so on”.

Almost every failure in this lesson would have been a line in such a record. The 79.6% in a composition field. The bad-words list in a processing field. The two darkly pigmented subjects out of ten in a collection field. The list below is that proposal at working length.

  • Purpose: why the dataset was created and which decisions it is intended to support.
  • Composition: units, time range, locations, languages, major slices, and outcome prevalence — the field where 79.6% lighter-skinned would have been written down.
  • Collection: source systems, sampling rules, participation, filters, and known missing channels.
  • Labeling: target definition, maturity window, reviewers, disagreement, and quality checks.
  • Processing: deduplication, joins, exclusions, transformations, and version identifiers — where a bad-words filter belongs.
  • Limits: unsupported uses, coverage gaps, privacy constraints, and expected refresh schedule.

Key idea

Narrowing the claim is a valid engineering result

When evidence is strong only for a subset — English messages from two products, say — the honest response may be to launch only there. Expanding the claim without data is not ambition. It is unsupported extrapolation. Six object-recognition systems were not wrong about household objects. They were right about household objects in the United States, and 15 points worse in Somalia or Burkina Faso. No sentence in their documentation drew that line.

A staged scope can create safer feedback and reveal what additional collection is genuinely needed.

The valid deployment boundary should follow the evidence, not the marketing plan.

Key takeaways