Research
FairFinGAN: Fairness-aware Synthetic Financial Data Generation
Overview Research area: Fairness-aware machine learning and synthetic tabular data generation, applied to financial datasets. Technical level: Advanced. The work builds on Wasserstein GAN training, ad

- arXiv
- 2603.05327
- Published
- 2026-03-05
- Authors
- Tai Le Quy, Dung Nguyen Tuan, Trung Nguyen Thanh, Duy Tran Cong, Huyen Giang Thi Thu, Frank Hopfgartner
AI summary
Overview
- Research area: Fairness-aware machine learning and synthetic tabular data generation, applied to financial datasets.
- Technical level: Advanced. The work builds on Wasserstein GAN training, adversarial optimization, Gumbel-Softmax, and group fairness definitions.
- Scope: The paper introduces FairFinGAN, a WGAN-based generator that injects statistical parity and equalized odds constraints into training so that synthetic financial tabular data is less biased while remaining useful for downstream classifiers, and evaluates it on five real-world financial datasets against CTGAN and TabFairGAN.
What This Paper Is About
Financial records are hard to share because of privacy and ownership restrictions, so synthetic data is often proposed as a substitute. However, generative models can reproduce or amplify the historical bias already present in the original data. This paper's goal is to generate synthetic financial tabular data that achieves statistical parity with respect to a protected attribute (such as sex, gender, race, or age) while still preserving enough utility for downstream predictive tasks.
Key Contributions
- FairFinGAN framework: A WGAN-based framework for fairness-aware synthetic financial data generation, inspired by the TabFairGAN approach.
- Two-phase fairness training strategy: Fairness constraints, specifically statistical parity and equalized odds, are folded into the GAN objective using a multi-layer perceptron classifier evaluated on generated samples, so bias is mitigated at the dataset level rather than the classification level.
- Two model variants: FairFinGAN-SP scores fairness with statistical parity, and FairFinGAN-EOd scores it with equalized odds.
- Extensive evaluation: Experiments on five real-world financial datasets, compared with CTGAN and TabFairGAN, measuring accuracy, balanced accuracy, and seven fairness measures (SP, EO, EOd, PP, PE, TE, ABROCA), with source code released at https://github.com/tailequy/FairFinGAN.
Main Findings
- Preserves fairness at the dataset level: Statistical parity was computed on the original and synthetic datasets. The reported statement is that FairFinGAN achieves the best or second-best values in most datasets with various protected attributes. For example, on Adult with protected attribute Gender, the original SP is 0.1989, CTGAN 0.1474, TabFairGan 0.0273, FairFinGAN-SP 0.1371, and FairFinGAN-EOd 0.1989.
- Trade-off rather than dominance on Adult: On Adult (Gender), FairFinGAN variants obtain accuracy and balanced accuracy slightly lower than CTGAN but higher than TabFairGAN, with generally better fairness (SP, ABROCA) than CTGAN, though not as extreme as TabFairGAN, which shows strong fairness but poor predictive performance. FairFinGAN-SP is slightly better for the LR classifier, while FairFinGAN-EOd is comparative in kNN.
- Mixed results on Adult (Race): Both variants show mixed results, but some measures (SP, EO, EOd, PP) improve with FairFinGAN-EOd for kNN and DT, and the paper reports a better fairness-accuracy trade-off than TabFairGAN.
- Strong results on Credit card (Sex): FairFinGAN-SP achieves the highest accuracy in most predictive models, and FairFinGAN-EOd often attains the best or second-best values for fairness metrics such as EO and EOd, particularly for DT and LR. TabFairGan performs well with the MLP classifier but has a missing value in TE.
- Credit card (Age): FairFinGAN-EOd achieves the best or second-best fairness measures (EO, EOd, PP, ABROCA) with MLP and LR, while TabFairGan shows strong accuracy for MLP and DT.
- Credit scoring: With protected attribute Sex, FairFinGAN-EOd achieves the best or second-best fairness results (SP, EO, EOd, PP, TE) while maintaining competitive accuracy across all predictive models. With protected attribute Age, FairFinGAN-SP shows balanced performance with high accuracy and improved SP in several classifiers (MLP and LR). TabFairGan performs well in accuracy but tends to exhibit weaker fairness across most measures.
- Dutch census: FairFinGAN-EOd achieves the highest accuracy across all classifiers while maintaining good fairness (EO, EOd, PP). FairFinGAN-SP reduces accuracy but yields the smallest deviations on fairness measures, especially SP, across all classifiers. CTGAN and TabFairGAN often introduce notable drops in accuracy.
- German credit: FairFinGAN-SP produces good fairness results (SP, EO, EOd, ABROCA), especially for LR and DT, sometimes with reduced accuracy. FairFinGAN-EOd improves fairness in MLP and DT. CTGAN often yields the highest accuracy and fairness (EO, EOd) across all classifiers.
- Model dependence: The variation in fairness gains between predictive models suggests the effectiveness of the synthetic data approach depends on the underlying learning mechanism, underscoring the need for model-specific evaluation.
- Reported caveats: TE can be unbounded because the number of false positives can be zero, and several result cells in the tables are reported as nan.
Methodology in Plain English
FairFinGAN trains in two phases. In phase one, a generator and a critic play the standard adversarial game of a Wasserstein GAN: the generator tries to make samples look real, and the critic is trained with a gradient penalty (weighted by lambda_pen) to score real versus generated batches. In phase two, the last part of training, a multilayer perceptron classifier that had already been trained on the real dataset is applied to the freshly generated samples. Its fairness score, either statistical parity or equalized odds, is computed on the predicted soft labels (made differentiable through Gumbel-Softmax) and added to the generator's loss, weighted by lambda_fair, so the generator is pushed toward producing data whose predicted outcomes do not differ between protected groups. The generator architecture uses fully connected ReLU layers with Gumbel-Softmax (tau = 0.2) for categorical outputs, the critic uses LeakyReLU layers, and the classifier is a two-hidden-layer MLP. Training runs for 200 total epochs, with the final 50 epochs in fairness mode, using a batch size of 256, a learning rate of 0.0002 for the generator and critic, 0.0001 for the fairness-aware generator in the second phase, lambda_fair = 0.5, lambda_pen = 10, and n_critic = 4. Evaluation uses four classifiers (Logistic Regression, Decision Tree, k-nearest neighbors, and MLP) with 5-fold cross-validation, comparing against CTGAN and TabFairGAN.
Why This Matters
- Impact on research: The paper moves fairness intervention upstream, into the data generation process, rather than only post-processing model outputs. It shows that fairness can be traded against utility in a controllable way and that the result depends heavily on which downstream classifier is used, which argues for model-specific fairness evaluation.
- Real-world applications:
- Lending and loan approval, where historical bias against protected groups can be embedded in training data.
- Credit scoring and risk assessment, where fairer synthetic data could support more equitable scoring models.
- Regulatory compliance work in financial institutions that need to demonstrate bias mitigation while sharing or publishing data substitutes.
- Financial data sharing and research collaboration, where privacy restrictions block access to original records.
- Industry relevance: The authors state that fairness-aware data generation can help reduce historical bias, promote more equitable risk assessment, and better align automated decisions with regulatory requirements while preserving predictive performance. Five of the datasets are drawn from finance and fairness-aware ML benchmarks, which makes the setup directly comparable to production-style tabular problems.
Future Directions
- Multiple protected attributes: Extending FairFinGAN to handle more than one protected attribute simultaneously, since the current design targets a single protected attribute per evaluation.
- Other domains: Applying the framework beyond finance, specifically to healthcare and education.
- Advanced fairness metrics: Investigating more advanced fairness metrics than the seven used here.
- Differential privacy: Incorporating differential privacy to improve the reliability and applicability of the generated data.
Target Audience
Researchers and practitioners working on fairness-aware machine learning, synthetic tabular data generation, and trustworthy AI in finance. It is also relevant to data scientists and compliance teams in banking and credit who need to produce shareable, less-biased data substitutes, and to readers already comfortable with GAN training and group fairness definitions.
Authors’ abstract
Financial datasets often suffer from bias that can lead to unfair decision-making in automated systems. In this work, we propose FairFinGAN, a WGAN-based framework designed to generate synthetic financial data while mitigating bias with respect to the protected attribute. Our approach incorporates fairness constraints directly into the training process through a classifier, ensuring that the synthetic data is both fair and preserves utility for downstream predictive tasks. We evaluate our proposed model on five real-world financial datasets and compare it with existing GAN-based data generation methods. Experimental results show that our approach achieves superior fairness metrics without significant loss in data utility, demonstrating its potential as a tool for bias-aware data generation in financial applications.