Research
Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial Inference
Overview Research area: fair representation learning (FRL), group fairness, and high-confidence statistical guarantees for machine learning. Technical level: Advanced. Scope: This paper introduces FRG

- arXiv
- 2510.21017
- Published
- 2025-10-23
- Authors
- Yuhong Luo, Austin Hoag, Xintong Wang, Philip S. Thomas, Przemyslaw A. Grabowicz
AI summary
Overview
Research area: fair representation learning (FRL), group fairness, and high-confidence statistical guarantees for machine learning. Technical level: Advanced. Scope: This paper introduces FRG, a framework for learning data representations that are guaranteed, with user-controllable high confidence, to keep downstream demographic disparity below a user-defined threshold.
What This Paper Is About
Fair representation learning aims to produce representations that are fair across many downstream tasks and models, rather than only for one specific task. Prior methods often estimate unfairness upper bounds from training or validation data, but these estimates may not hold on unseen test data. This paper formalizes the problem of learning representations with high-confidence fairness guarantees and proposes FRG to guarantee that downstream demographic parity disparity stays within a user-specified threshold with controllable high probability.
Key Contributions
- The paper formally introduces the task of learning representations that achieve high-confidence fairness, where demographic disparity in every downstream prediction is bounded by a user-defined error threshold ε with controllable high probability. FRG is described as the first work to guarantee, with high confidence, that output representation models are fair according to user-specified thresholds of unfairness and confidence levels.
- FRG is built from three components: candidate selection, adversarial inference, and a fairness test. Candidate selection proposes a representation model likely to pass the fairness test; adversarial inference trains an adversary to predict sensitive attributes from representations; the fairness test constructs a high-confidence upper bound on worst-case downstream Δ_DP.
- The paper proves a theoretical mapping between Δ_DP and the absolute covariance between the sensitive attribute and downstream predictions when both are binary (Theorem 5.2), showing that an optimal adversary can be approximated by maximizing absolute covariance. The fairness test then uses confidence intervals from standard statistical tools such as Student’s t-test and Hoeffding’s inequality.
- The paper empirically evaluates FRG on three real-world datasets against six state-of-the-art fair representation learning methods. Source code is available at https://github.com/JamesLuoyh/FRG.
Main Findings
-
High-confidence fairness is maintained by FRG: FRG and FRG_supervised can maintain Δ_DP ≤ ε with sufficiently high probability of at least 0.9. Most baseline methods cannot consistently satisfy ε-fairness with high probability across all datasets.
-
Baselines fail more often with smaller ε: For baseline methods, a smaller ε causes a larger probability of failing the constraint. Baselines also tend to fail on adversarial tasks, in contrast to FRG.
-
FRG preserves predictive performance: FRG can match or outperform baselines in AUC. Compared with baselines that achieve Δ_DP ≤ ε with high probability, such as ICVAE for Adult and Health and FARE, FRG tends to yield higher AUC, especially when ε is small.
-
FRG rarely returns No Solution Found: FRG keeps a high probability of returning a solution, at least 0.9 over all datasets, which demonstrates the effectiveness of candidate selection.
-
FARE comparison: FARE can also achieve a high probability of satisfying the fairness constraint across all datasets. However, FARE does not support user-defined ε values, and its certificates are loose and often several times larger than the desired ε. Even with hyperparameter search, FARE does not certify fairness with enough granularity, giving the same certificates to multiple ε values.
-
Dataset-specific evaluations: On Adult, the target label is income and the sensitive attribute is gender, with ε in {0.04, 0.08, 0.12, 0.16} and δ fixed at 0.1. On Health, the sensitive attribute is gender, the original target task predicts Charlson Index, and transfer learning is shown for predicting age. On Income, the target label is income and the sensitive attribute is marital status with 5 classes, with ε in {0.16, 0.24, 0.32, 0.40, 0.48}.
Methodology in Plain English
FRG splits the available data into two disjoint sets: D_c for candidate selection and D_f for the fairness test. Candidate selection searches over representation models using D_c, optimizing a constrained objective that balances reconstruction or prediction quality against an inflated high-confidence upper bound on unfairness. The inflation factor α is a hyperparameter with α ≥ 1, used to reduce overfitting to D_c and avoid returning No Solution Found too often.
Separately, adversarial inference trains an adversary to predict the sensitive attribute from the learned representations. The paper shows that the worst-case demographic parity disparity equals the absolute covariance between the sensitive attribute and the prediction divided by the variance of the sensitive attribute, so the optimal adversary is approximated by maximizing absolute covariance. The adversary is trained with D_c.
The fairness test then uses D_f and the trained adversary to estimate positive prediction rates conditioned on sensitive attribute values. It constructs a 1 − δ confidence interval on the difference between those rates, using procedures such as Student’s t-test. If the resulting confidence upper bound on g_ε is at most zero, the candidate representation model is considered ε-fair with confidence 1 − δ, and FRG returns it. Otherwise, FRG returns No Solution Found. FRG does not search for and test another model after a failure, because doing so would create a multiple comparisons problem.
Why This Matters
This work matters because it moves fair representation learning from empirical estimates toward explicit, controllable statistical guarantees. It lets a data producer specify both an unfairness threshold ε and a confidence level, then provides representations that are intended to be safe across arbitrary downstream tasks and models, including adversarial ones. This is important for research on trustworthy machine learning, and for legal and policy settings where group fairness is scrutinized, such as the New York City Local Law 144 on Automated Employment Decision Tools and the EEOC’s rule of 80% hiring rates across sensitive groups.
Real-world applications include:
- Loan underwriting, where unfair predictions can affect access to credit.
- Hiring, where automated tools may disadvantage demographic groups.
- Criminal sentencing, where algorithmic bias can have severe consequences.
- General representation learning for text or images, where one data producer creates representations used by many downstream data consumers.
Industry relevance: organizations that generate general representations can use FRG to provide downstream teams with representations that carry controllable fairness guarantees, reducing the need for each downstream consumer to independently enforce fairness. This is especially relevant when downstream tasks are unknown or unlabeled at representation-learning time.
Future Directions
- Extend the guarantees beyond binary classification and demographic parity to Equal Opportunity and Equalized Odds, as discussed in Appendix C, and to non-binary sensitive attributes and regression settings discussed in Appendix D and Appendix A.
- Improve the approximation of the optimal adversary and avoid relying on loose theoretical upper bounds. The paper notes that previous upper bounds are typically loose and impractical for establishing guarantees while preserving utility, with a non-trivial gap demonstrated in Appendix Figures 12 and 13.
- Address the multiple comparisons problem if users want to search and test multiple candidate representation models after a failed fairness test. Currently, FRG returns No Solution Found rather than continuing to search.
- Further develop and compare certification methods such as FARE, which can provide high-confidence certificates but does not support user-defined ε and gives the same certificates to multiple ε values.
Target Audience
This paper benefits researchers and practitioners in fair machine learning, representation learning, and trustworthy ML, especially those working on statistical guarantees, group fairness, and adversarial inference. It is also relevant to data producers who create general representations for downstream use, industry teams deploying models in high-stakes domains such as finance, hiring, and healthcare, and regulators or policy analysts interested in enforceable fairness criteria.
Authors’ abstract
Representation learning is increasingly applied to generate representations that generalize well across multiple downstream tasks. Ensuring fairness guarantees in representation learning is crucial to prevent unfairness toward specific demographic groups in downstream tasks. In this work, we formally introduce the task of learning representations that achieve high-confidence fairness. We aim to guarantee that demographic disparity in every downstream prediction remains bounded by a *user-defined* error threshold $ε$, with *controllable* high probability. To this end, we propose the ***F**air **R**epresentation learning with high-confidence **G**uarantees (FRG)* framework, which provides these high-confidence fairness guarantees by leveraging an optimized adversarial model. We empirically evaluate FRG on three real-world datasets, comparing its performance to six state-of-the-art fair representation learning methods. Our results demonstrate that FRG consistently bounds unfairness across a range of downstream models and tasks.