Research
SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
Overview Research area: Privacy-preserving machine learning, specifically the evaluation of synthetic data generation (SDG) for health data — privacy metrics, membership inference, identity disclosure

- arXiv
- 2601.12124
- Published
- 2026-01-17
- Authors
- Bing Hu, Yixin Li, Asma Bahamyirou, Helen Chen
AI summary
Overview
- Research area: Privacy-preserving machine learning, specifically the evaluation of synthetic data generation (SDG) for health data — privacy metrics, membership inference, identity disclosure, differential privacy, and synthetic data quality.
- Technical level: Intermediate. The paper assumes familiarity with generative models (GANs, VAEs, copulas), differential privacy, and standard privacy-attack concepts, though the framework itself is described procedurally.
- Scope in one sentence: The paper introduces SynQP, an open framework for benchmarking the quality and privacy risk of synthetic data using simulated pseudo-identifiable data, and uses it to benchmark CTGAN, TVAE, and GaussianCopula (with and without differential privacy) while proposing new synthetic-data identity disclosure and membership inference metrics.
What This Paper Is About
Synthetic data is promoted as a privacy-enhancing technology for sharing sensitive health information, but there is no open, comparable way to measure how much privacy risk a synthetic dataset actually carries. The core obstacle is that benchmarking privacy metrics requires data with identifying information, which researchers usually cannot access or redistribute. SynQP addresses this by simulating a pseudo-identifiable population from already de-identified real data, so that privacy risks can be evaluated and compared without exposing anyone's real information.
Key Contributions
- An open benchmark framework (SynQP) that standardizes the evaluation of privacy risk for synthetic data generation models by constructing a simulated pseudo-identifiable population linked to non-identifiable real data.
- A new identity disclosure risk definition (SD-IDR) that replaces exact-match cardinality with a tolerance-based match indicator, allowing records that differ only by small numerical variations to count as matches.
- A new synthetic data membership inference attack metric (SD-MIA) defined as the difference between identity disclosure risk against synthetic data and against a random sample of the real distribution, which indicates whether a generator is overfitting, underfitting, or behaving like chance.
- An empirical demonstration applying SynQP and SD-IDR to CTGAN, TVAE, and GaussianCopula, with and without differential privacy, on a simulated population built from a de-identified diabetes dataset.
Main Findings
- Differential privacy lowers measured privacy risk: Privacy assessments (Table II) show that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold set by Health Canada and the European Medical Agency.
- Exact-match IDR can underestimate risk: As the variational budget increases from 0 (exact matches only) through 1, 2, and 3, SD-IDR rises for every model, so the authors argue that multiple variational budgets, including 0, should be evaluated for a holistic view of identity disclosure risk.
- Fidelity varies sharply by model: Average Hellinger distances without DP are 0.18 for CTGAN, 0.06 for GaussianCopula, and 0.34 for TVAE (with the paper's text reporting DP-augmented averages of 0.27, 0.24, and 0.34 respectively, while Table 1 lists a TVAE-DP average of 0.51). Lower Hellinger distance indicates a closer match between real and synthetic distributions.
- Utility drops when DP is added: Machine learning efficiency (AUC) scores without DP are 0.97 for CTGAN, 0.99 for GaussianCopula, and 0.99 for TVAE; with DP they fall to 0.47, 0.94, and 0.87 respectively, illustrating a trade-off between utility, fidelity, and privacy.
- DP visibly degrades the generated distribution: Figure 2 shows that differential privacy greatly degrades the overall distribution of generated synthetic data compared to the original training data for age and BMI.
- SD-MIA separates overfitting from underfitting: TVAE's SD-MIA increases as the variational budget increases, whereas CTGAN and GaussianCopula SD-MIA decrease toward negative values. The authors interpret this as TVAE grossly overfitting certain training rows, while CTGAN and GaussianCopula may still be underfit and could undergo additional training.
- The framework extends earlier work: SynQP was previously applied to evaluate CTGAN, where IDR was found to potentially underestimate re-identification risk compared to SD-IDR; this study extends that result to three SDG models and their DP variants.
Methodology in Plain English
The framework works in three stages.
First, it builds a simulated population. Quasi-identifiers — pieces of information that are not unique on their own but can combine to identify someone — are generated by sampling from real distributions. Age is sampled with inverse transform sampling from census distributions and spans 0 to 99; gender is conditionally sampled given age for men+ or women+; marital status, occupation, ethnicity, and address are randomly sampled from lists of 7 marital statuses, 1154 occupations, and 250 ethnicities, with addresses generated using an available Python package. The authors note they do not currently model correlations between age and gender and the other attributes.
Second, real non-identifiable data is linked to each simulated row. For the demonstration, the authors use diabetes data from the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) and publicly available BMI data. Height, weight, and BMI are conditionally sampled given gender using the BMI distributions, and the diabetes columns are infilled by finding the nearest neighbour in the diabetes dataset using age and BMI; k-nearest neighbours is mentioned as an alternative.
Third, synthetic data is generated and evaluated. The authors simulated a population of 10,000 rows for the diabetes dataset, sampled 7,000 rows as the training set with quasi-identifier columns age, gender, and marital status and data columns BMI and number of pregnancies, and kept the remaining 3,000 rows as holdout. CTGAN, TVAE, and GaussianCopula were trained with default hyperparameters from the Python library Synthetic Data Vault to generate 10,000 synthetic rows. Differential privacy was applied as a local mechanism injecting Laplace noise into examples before training, using a privacy budget ε in [0,1] with ε levels of 0 and 0.8, where the authors note 0.8 is quite large but sufficient for demonstration. Evaluation used Hellinger distance for fidelity, machine learning efficiency with a logistic regression measured by AUC for utility, and the proposed SD-IDR and SD-MIA for privacy. For MIA, 3,000 rows from the simulated population were treated as data known to an attacker.
Why This Matters
Synthetic data is often justified on privacy grounds, yet privacy evaluation is frequently omitted from SDG publications because comparable benchmarks do not exist. SynQP supplies a shareable, open benchmark that lets models be compared on privacy as well as quality, and its metrics respond to the probabilistic nature of generative models rather than treating near-matches as non-matches. More broadly, the authors position SynQP as a translational tool between vague regulatory language and testable technical requirements.
Real-world applications:
- Health data holders and administrative data custodians can liberalize data access by generating and releasing simulated pseudo-identifiable datasets for benchmarking rather than distributing real patient records.
- Regulatory compliance demonstration — organizations can quantify utility and privacy risk against thresholds such as the 0.09 identity disclosure limit used by Health Canada and the European Medical Agency.
- Model selection and procurement — teams can construct an evaluation matrix of models and privacy budgets to identify combinations that satisfy both quality and privacy requirements before deployment.
- Research enablement — open simulated datasets allow methods to be compared across institutions without lengthy approvals or data-sharing agreements.
Industry relevance: the work targets pharmaceutical, health-system, and clinical research organizations that want to use real-world data such as electronic medical records and electronic health records but face access barriers, as well as vendors of privacy-enhancing technologies who need defensible, standardized evidence that their generators meet privacy expectations.
Future Directions
- Further develop the SynQP framework and apply it to additional SDG models and datasets beyond CTGAN, TVAE, and GaussianCopula.
- Publish an extensive open dataset specifically for future SDG benchmarking, as the authors state they intend to do.
- Incorporate additional quasi-identifiers and reconsider attribute correlations, since the current simulation does not model correlations between age and gender and occupation, marital status, address, or ethnicity.
- Explore alternative sampling methodologies for linking real data, such as k-nearest neighbours, and adapt SD-MIA to additional adversarial sampling scenarios.
Target Audience
Researchers and engineers building or auditing synthetic data generators, particularly in health and clinical settings; privacy and statistical disclosure researchers interested in identity disclosure and membership inference metrics; data governance, policy, and regulatory staff who need to translate privacy requirements into measurable technical criteria; and data custodians deciding whether synthetic data can be safely released.
Authors’ abstract
The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQP to benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. % In our quality evaluations, non-private models achieved near-perfect machine-learning efficacy \(\ge0.97\). Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold. Code available at https://github.com/CAN-SYNH/SynQP