Skip to content
AI.info

Recommender systems

Fairness, Exposure, and User Control

Analyze user-side and provider-side fairness, exposure allocation, subgroup performance, contestability, and personalization controls.

By the end you can

Example

Identical targeting, and the platform split the audience anyway

Same targeting, same bidding strategy, same budget, all running at the same time. Everything the advertiser controls, held fixed. A 2019 field study bought real employment and housing ads on Facebook that way. That left the platform's own delivery stage as the only thing free to vary. It varied a great deal.

“In the most extreme cases, our ads for jobs in the lumber industry reach an audience that is 72% white and 90% male, our ads for cashier positions in supermarkets reach an 85% female audience, and our ads for positions in taxi companies reach a 75% Black audience, even though the targeted audience specified by us as an advertiser is identical for all three.” — Muhammad Ali and colleagues, 2019.

Housing ads run under the same conditions split the same way, from audiences over 72% Black users down to as little as 51% Black users.

Now look at where the disparity sits. An advertiser auditing its own campaign settings would find them identical across the three job ads, because they were identical. The split happened after the settings, in a stage the buyer does not write and the recipient never sees. An audit that stops at the targeting parameters inspects the one part of the pipeline that was demonstrably clean.

  • Identical inputs: same targeting, same bidding strategy, same budget, same time — everything the advertiser controls was held fixed across the three job ads.
  • Provider-side exposure: the delivered audiences were 72% white and 90% male for lumber-industry jobs, 85% female for supermarket cashier positions, and 75% Black for taxi-company positions.
  • The same effect in housing: under the same conditions, housing ads landed on audiences ranging from over 72% Black users down to as little as 51% Black users.
  • Where the allocation lives: the split is produced by the delivery stage, so an audit that stops at the targeting parameters examines the one part of the pipeline that was demonstrably clean.
  • Contestability: neither the buyer who set identical parameters nor the person who was never shown the ad can inspect or correct the decision that separated them.

Example

Fairness failures, and what a group average hides

Protected-attribute removal and decorative controls are two forms of the same substitution: the appearance of a remedy in place of one. Drop the attribute and the proxies stay where they are, while the audit gets harder. Ship a “show less” button that barely alters the policy and the user gets a lever attached to nothing.

The group average has a measured case against it. Gender Shades audited three commercial gender classifiers in 2018. Its abstract: “We evaluate 3 commercial gender classification systems using our dataset and show that darker-skinned females are the most misclassified group (with error rates of up to 34.7%). The maximum error rate for lighter-skinned males is 0.8%.” — Joy Buolamwini and Timnit Gebru, 2018.

34.7% against 0.8% is a gap that appears only when sex and skin type are crossed. Report accuracy by sex alone, or by skin type alone, and it flattens into something a launch review would wave through.

The benchmarks were part of the problem. The two audited there, IJB-A and Adience, were 79.6% and 86.2% lighter-skinned. A system could post a strong aggregate score on them and still fail the group they barely sample.

  • Protected-attribute removal: proxies remain while auditing becomes harder — the Facebook delivery split was measured without any advertiser ever supplying a demographic attribute.
  • Exposure without relevance: a parity target ignores user need, provider quality, or eligibility.
  • Average subgroup metric: intersectional and low-volume harms disappear — 34.7% error for darker-skinned females sits behind an aggregate that also contains 0.8% for lighter-skinned males.
  • Benchmark composition: a test set that is 79.6% or 86.2% lighter-skinned rewards the aggregate and conceals the slice it barely samples.
  • Decorative controls: “show less” or opt-out settings barely alter the policy.

Visual

A fairness assessment across the recommendation pipeline

Ask the same question at every stage, because the answer keeps changing. Data and exposure: who was ever visible, who interacted, and which histories are missing — the point where a benchmark that is 79.6% lighter-skinned decides what "accurate" will mean. Candidate generation: which groups or providers can enter the set at all. Ranking and position: how attention is allocated within the relevant candidates. Outcome and burden: who benefits, who is harmed, and who must do the extra work; in the Facebook study the burden fell on people who never learned that a job ad existed. Control and remedy: whether a person can inspect, correct, opt out, appeal, or choose another mode.

That last stage can be measured against a written standard rather than a preference. The Digital Services Act makes the non-profiled option compulsory: “In addition to the requirements set out in Article 27, providers of very large online platforms and of very large online search engines that use recommender systems shall provide at least one option for each of their recommender systems which is not based on profiling as defined in Article 4, point (4), of Regulation (EU) 2016/679.” — Article 38, Regulation (EU) 2022/2065.

Article 27(3) then settles where the control has to live. The switch between available options must be directly and easily accessible from the interface section where the information is being prioritised. Between them, the two articles give the decorative-control failure something concrete to fail against. A control buried three menus away from the feed it governs does not satisfy Article 27(3). A personalisation-only recommender does not satisfy Article 38 at all.

FigureTimeline · 5 stops
  1. Data and exposure

    Who was visible, who interacted, and which histories are missing?

  2. Candidate generation

    Which groups or providers can enter the set at all?

  3. Ranking and position

    How is attention allocated within relevant candidates?

  4. Outcome and burden

    Who benefits, who is harmed, and who must do extra work?

  5. Control and remedy

    Can users inspect, correct, opt out, appeal, or choose another mode?

Fairness in recommendation concerns users, providers, and allocation over time

User fairness can involve relevance, error, opportunity, burden, accessibility, or control. Provider fairness can involve position-weighted exposure relative to relevance, merit, need, or another justified allocation rule.

No fairness metric defines justice by itself. Which target is the right one depends on the surface, on who stands in what relationship to whom, on the legal context the system runs in, on which candidates were eligible at all, and on the harm done by showing something as much as by leaving it out.

The two measured cases above pick out two different targets, and neither would have caught the other. The Facebook study is about who received an ad the advertiser wanted delivered to everyone. Gender Shades is about error rates on a group the benchmark barely contained.

Choosing the target is the substantive decision, and letting the most convenient metric make it settles who is owed what without ever arguing the point.

Steps

Conduct a recommendation fairness review

Harms come first in a fairness review, together with the people who carry them: recipients, providers, workers, and communities that never see the interface. Step 2 locates the allocation points — data, eligibility, retrieval, ranking, layout, and feedback. That is what the Facebook design did, holding the advertiser's parameters fixed until only delivery was left.

Step 3, select justified metrics. "Context-justified" is not a placeholder. A regulator has written one down. New York City's rule implementing Local Law 144 of 2021 took effect on 5 July 2023. It defines the "impact ratio" as the selection rate for a category divided by the selection rate of the most selected category. And it refuses to take that as a group average: “Clarifying that the required “impact ratio” must be calculated separately to compare sex categories, race/ethnicity categories, and intersectional categories;” — New York City Department of Consumer and Worker Protection. The intersectional slice is required, not encouraged.

Step 4, test interventions. LinkedIn published a fairness-aware re-ranking framework in 2019 and A/B tested it on a live hiring surface: “Our approach resulted in tremendous improvement in the fairness metrics (nearly three fold increase in the number of search queries with representative results) without affecting the business metrics, which paved the way for deployment to 100% of LinkedIn Recruiter users worldwide.” — abstract of LinkedIn's fairness-aware ranking paper, 2019. Close to a threefold increase in the number of search queries returning representative results. No measured loss on business metrics. Then rollout to 100% of LinkedIn Recruiter users worldwide, across a member base of more than 630 million. That is the shape a tested intervention reports in: an effect size, a stated utility outcome, and a named surface.

Step 5 is remedy — correction, appeal, reset, and monitoring that somebody owns by name.

FigureProcess · 5 steps
  1. 1. Map stakeholders and harms

    Include recipients, providers, workers, and affected communities.

  2. 2. Locate allocation points

    Audit data, eligibility, retrieval, ranking, layout, and feedback.

  3. 3. Select justified metrics

    Tie relevance, exposure, opportunity, burden, and control to the context.

  4. 4. Test interventions

    Compare data, model, re-ranking, product, and institutional changes.

  5. 5. Provide remedy

    Implement effective correction, appeal, reset, and monitoring.

Fairness evidence should survive policy changes

Report candidate recall, position-weighted exposure, user utility, outcome, complaints, and control effectiveness across relevant groups. Evaluate both absolute levels and changes from the previous policy. When provider exposure is constrained, document the allocation principle and test consumer utility, capacity, quality, and long-term supply. A mathematically balanced ranking can still be socially or commercially unjustified.

Control effectiveness is measurable, and the measurement is unkind. A 2024 study checked whether New York City's audit-and-notice remedy had moved anything: “In this study, 155 student investigators recorded 391 employers’ compliance with LL 144 and the user experience for prospective job applicants. Among these employers, 18 posted audit reports and 13 posted transparency notices.” — Wright and colleagues, 2024. Of the impact ratios that were posted, only nine of 386 fell below the 0.8 four-fifths threshold. So 18 of 391 employers acted, and the audits they did post almost never report a failing number. That is a disclosure regime measured rather than assumed.

The baseline requirement has a worked example too. The experiment ran on Twitter, and it began by holding a control group off the algorithm: “We provide quantitative evidence from a long-running, massive-scale randomized experiment on the Twitter platform that committed a randomized control group including nearly 2M daily active accounts to a reverse-chronological content feed free of algorithmic personalization.” — Huszár and colleagues, PNAS, 2022. Every amplification figure was then read against that held-out feed. The comparison covered 3,634 legislator accounts in seven countries and 6.2 million news articles. In six of the seven countries the mainstream political right was more amplified than the mainstream political left. There was no evidence that far-left or far-right groups were amplified more than moderates. Without the control group, the same exposure counts would have described a state and attributed it to nothing.

Without the previous policy as a baseline, a group-level exposure number describes a state and settles nothing about the change that produced it.

Case

Fairness written as an allocation of exposure

Fairness constraints can be stated as allocations of exposure rather than as properties of a score. Singh and Joachims gave that a formal shape in 2018. Demographic parity, disparate treatment and disparate impact can all be written that way, and the ranking is then chosen to maximise utility subject to those constraints.

The useful part is that exposure is measurable where fairness is arguable. The Facebook delivery study relied on that property to make its case. New York City's impact ratio borrows it to make an audit computable.

Key idea

The fairness gate

Deploy a fairness intervention only when the affected allocation, target principle, relevance tradeoff, and remedy path are explicit.

One instrument names all four. In United States v. Meta Platforms, filed on 21 June 2022, the Justice Department settled Fair Housing Act claims over ad delivery: “Meta will build a Variance Reduction System (“VRS”) for Housing Advertisements to reduce variances in Ad Impressions between Eligible Audiences and Actual Audiences, as those terms are defined in the proposed Settlement Agreement, for sex and estimated race/ethnicity, that the United States alleges are introduced by Meta’s ad delivery system.” — the United States' memorandum in support of the proposed settlement.

Read it against the gate. The affected allocation is ad impressions by sex and estimated race/ethnicity. The target principle is the gap between Eligible Audience and Actual Audience. The relevance tradeoff is bounded by what the settlement removes: Meta was to stop delivering housing ads targeted with the Special Ad Audience tool, and to take that tool and the Lookalike Audience tool out of its housing ad flows by 31 December 2022. The remedy path is an independent third-party reviewer verifying agreed compliance metrics. Disputes go to the court, and the maximum civil penalty under the Act is attached, $115,054.

Meta launched the VRS for US housing ads on 9 January 2023 and announced its extension to US employment and credit ads on 26 October 2023. All four items were named in writing before the system existed.

An intervention shipped without a named relevance tradeoff and a named remedy path still moves the allocation — it just moves it where nobody agreed to look.

Key takeaways