Skip to content
AI.info

Research

Harnessing Synthetic Data from Generative AI for Statistical Inference

Overview Research area: Statistics and machine learning — specifically statistical inference, uncertainty quantification, and robustness when synthetic data produced by deep generative models (GANs, V

Harnessing Synthetic Data from Generative AI for Statistical Inference
arXiv
2603.05396
Published
2026-03-05
Authors
Ahmad Abdel-Azim, Ruoyu Wang, Xihong Lin

AI summary

Overview

Research area: Statistics and machine learning — specifically statistical inference, uncertainty quantification, and robustness when synthetic data produced by deep generative models (GANs, VAEs, normalizing flows, autoregressive/LLM models, diffusion models) are used in downstream analysis.

Technical level: Intermediate to Advanced. The paper is a review article rather than a new empirical study. It introduces formal notation (observed data, target population, sampling distribution, access patterns) and cites methods across privacy, fairness, domain adaptation, and missing-data literatures.

Scope: A statistical review of when synthetic data can validly support downstream discovery, inference, and prediction — with emphasis on the conditions and assumptions that make synthetic data use principled, particularly when the generative model is misspecified.

Note on the provided content: the text supplied covers the abstract, Section 1 (Introduction), and Section 2 (Synthetic data generation, including Section 2.1 on motivations and Section 2.2 on generative model classes) up to a mid-sentence discussion of conditional generation with transformers. The sections on statistically principled paradigms for using synthetic data, worked examples, practical recommendations, and open problems are described in the abstract and introduction but their detailed contents are not present in the truncated text.

What This Paper Is About

The paper addresses a gap that has opened up as generative AI has become widely used: practitioners now produce and consume large amounts of synthetic data, but the statistical conditions under which that data can support valid inference are not well understood. The authors argue that generative modeling capability has outpaced understanding of when synthetic data can reliably support downstream statistical inference and scientific discovery, and that naively treating synthetic samples as if they were real observations can lead to invalid inference and bias. Their goal is to organize the motivations for synthesis, survey the model families that produce synthetic data, and delineate the assumptions and failure modes that govern downstream utility.

Key Contributions

  1. An organizing framework for why synthetic data are generated. The paper separates synthetic data use into five motivations — privacy-preserving release, data augmentation, fairness, domain transfer, and missing data/trajectory completion — and characterizes each by two axes: (i) the target sampling distribution Q from which synthetic records are drawn, and (ii) the "access pattern" describing how an analyst interacts with the real data and the synthetic data downstream.

  2. A statistical formalization of synthetic data generation. The paper writes the problem in terms of predictors X, labels Y, an observed dataset of n observations from an unknown training distribution P, a possibly different target population P_T, an estimated generative distribution P̂_θ, an induced sampling distribution Q, and a synthetic dataset S. Downstream goals target either a functional β(P_T) (e.g., a causal estimand or regression parameter) or a predictive object m_{P_T} (e.g., a risk score or conditional mean).

  3. A comparative survey of deep generative model classes framed by their statistical objects. Table 2 compares GANs, VAEs, normalizing flows, autoregressive/Transformer models, and diffusion/score-based models, identifying each class's statistical object, core idea, and strengths and limitations relevant to downstream use — including whether an explicit likelihood p_θ(·) is available or only an implicit sampler.

  4. Identification of characteristic failure modes and pitfalls. The review highlights model misspecification, attenuated uncertainty, generalization difficulties, mode collapse, posterior collapse, and "model collapse" under naive recursive training, alongside the observation that privacy constraints deliberately shift Q away from P̂_θ, making a privacy-utility tradeoff explicit.

Main Findings

  • Synthetic data use is expanding well beyond its original purpose. Synthetic data were first proposed as a privacy-preserving measure, but modern uses include improving fairness, augmenting dataset size to increase inferential power, creating data for new domains, and generating synthetic tasks for in-context learning. Synthetic data also permeate LLM training pipelines, where utility is fairly established in post-training stages such as instruction-tuning and alignment, but uses and effects in pre-training are poorly understood.

  • Naive use of synthetic data can invalidate inference. Treating synthetic data the same as observed data in a statistical analysis — for example, simply combining real and synthetic samples without calibration or weighting — can lead to invalid inference and bias.

  • Model misspecification is central, not incidental. Synthetic samples drawn from misspecified generative models can systematically misrepresent key features of the target distribution, so understanding how synthesis error and misspecification propagate through inferential and predictive workflows is critical for robustness.

  • Recursive training on synthetic data degrades distributions. Naive recursive training of LLMs solely on generated outputs from earlier generations has been shown to lead to "model collapse," where learned models progressively lose diversity and misrepresent the tails of the original data distribution.

  • Privacy protection deliberately distorts the sampling distribution. Under (ε, δ)-differential privacy, the mechanism M must satisfy P(M(O) ∈ B) ≤ e^ε P(M(O′) ∈ B) + δ for adjacent datasets differing by one individual record. This forces deliberate random perturbation, so Q is by design shifted away from P̂_θ. In the canonical private mean estimation problem, clipping sample values to a bounded range controls sensitivity but introduces a non-vanishing statistical bias that cannot be removed without violating differential privacy; weaker privacy guarantees reduce this bias but at a cost to fidelity.

  • Multiple-imputation framing propagates synthesis uncertainty by construction. In the classical multiple-imputation formulation, the target distribution is the posterior predictive mixture Q(z) = ∫ P̂_θ(z) p(θ | O) dθ, where each synthetic release draws θ^(m) ~ p(θ | O) and then samples from P̂_θ^(m). Downstream uncertainty then reflects both sampling variability and uncertainty in the generative model parameters.

  • Conditioning changes the target, not just the sample. In conditional augmentation for severe class imbalance or rare events, the target sampling law is set to a conditional such as Q(·) = P̂_θ(· | A = a) or P̂_θ(· | Y = 1), so synthetic mass is allocated to a designated region of the data space rather than approximating P marginally.

  • Fairness synthesis targets a constrained distribution that intentionally differs from P. The formulation is Q* ∈ arg min_{Q ∈ Q} D(Q, P̂_θ) subject to Fair(Q) ≥ τ, with a canonical example Fair(Q) = −|P(Ŷ = 1 | A = 0) − P(Ŷ = 1 | A = 1)|. Here synthetic data are not a privacy-preserving or power-enhancing surrogate for the real data, but a principled way to adjust the training distribution.

  • Domain transfer requires a different target entirely. In multi-site settings the observed training data follow P while the deployment target is P_T ≠ P; the sampling distribution aims at Q ≈ P_T using source observations and typically some target information. A common special case is covariate shift, where P_T(Y | X) = P(Y | X) but P_T(X) ≠ P(X).

  • Trajectory completion is a trajectory-level, not static, conditional problem. The object of interest is Q = P̂_θ(Z_miss | Z_obs, A), with forecasting as the special case where the observed index set is {1, …, t₀}; the paper notes this differs from other conditional synthesis motivations because the trajectory-level conditional must respect temporal dependence.

  • Generative model families trade off likelihood tractability against fidelity. GANs produce sharp samples but suffer training instability and mode collapse; VAEs offer principled probabilistic models and stable training but can produce blurrier samples and suffer posterior collapse; normalizing flows give exact likelihoods and invertible inference but face architectural constraints for tractability; autoregressive/Transformer models give tractable factorized likelihoods and strong conditional generation but sample slowly and often need huge data; diffusion/score-based models achieve state-of-the-art fidelity and diversity with stable training but are expensive during iterative refinement.

  • Tabular synthesis has modality-specific failure modes. Mixed variable types, skewed marginals, long-tailed continuous variables, and imbalanced categories make high-fidelity tabular generation challenging; failure to capture these dependencies produces synthetic artifacts that degrade downstream utility rather than expanding effective sample size.

Methodology in Plain English

This is a review and synthesis paper, so the "method" is conceptual organization rather than an experiment. The authors proceed in three steps.

First, they fix notation that lets them talk about all synthetic data uses in one language. Real data are drawn from an unknown training distribution P; a target population P_T may differ from it; the generative model produces an estimated distribution P̂_θ, which depends on an unknown parameter θ; synthetic data are then drawn from an induced sampling distribution Q, often Q = P̂_θ. Downstream tasks are functionals β(P_T) or predictive objects m_{P_T} of the target distribution.

Second, they separate the reasons for synthesizing data along two axes: what the target sampling distribution Q is intended to be, and what access pattern the analyst has to the real and synthetic data. This lets them distinguish settings where Q is meant as a plug-in approximation to P (privacy release, augmentation), where Q is deliberately conditional on an attribute or label (targeted augmentation, imputation, forecasting), where Q is shifted toward a different population (domain transfer), and where Q is constrained to satisfy an external criterion (privacy guarantee, fairness threshold). Each of the five motivations is then illustrated with named example methods and their access patterns.

Third, they survey the major deep generative model classes at the level of their statistical objects — what each model estimates, whether an explicit likelihood is available, and what its characteristic strengths and limitations are for downstream use. Architectural and training details are deliberately left out, with the authors pointing to existing comprehensive reviews for those. The attention mechanism is given as a concrete illustration of how Transformer-based autoregressive models capture long-range dependence, using the standard softmax(QKᵀ/√d_k)V form.

Throughout, the authors flag where statistical assumptions are needed for valid downstream use, such as covariate shift for importance-weighting approaches, and where the generative procedure itself imposes a distortion, such as the deliberate perturbation required by differential privacy.

Why This Matters

Impact on research. The paper reframes synthetic data as a statistical object with assumptions rather than a generic data source. It makes explicit that the validity of an analysis using synthetic data depends on the relationship between the sampling distribution Q and the target P_T, on the access pattern, and on whether the generative model is correctly specified. This gives statisticians and machine learning researchers a shared vocabulary for stating guarantees and for identifying where existing theory does not apply. It also flags that simple pooling of real and synthetic data is not a valid default.

Real-world applications (from the paper's own examples):

  • Privacy-preserving data release. Releasing synthetic records instead of individual-level records so external analysts can run analyses without accessing the original data, using frameworks such as multiple imputation, DP-SGD, PATE, PrivBayes, and PATE-GAN. The paper emphasizes that this comes with an unavoidable privacy-utility tradeoff, since privacy constraints intentionally push the synthetic distribution away from the data-generating distribution.

  • Medical and clinical research. The paper discusses augmenting a modest clinical cohort to train more stable predictors, oversampling rare adverse events in a cohort with Y ∈ {0, 1}, imputing missing clinical measurements, and "digital twin" formulations that condition on patient history and potential interventions to generate forward trajectories.

  • Multi-site biomedical prediction. Training on a source hospital dataset and deploying at a target hospital with different covariate distributions, measurement practices, or prevalence, using synthetic transfer samples, importance-weighted cross-validation, RadialGAN-style cross-dataset mappings, or optimal transport to align domains.

  • Fairness-aware lending and risk scoring. Using constrained synthesis so that a predictor for outcomes such as loan default meets fairness criteria with respect to a protected attribute (for example, demographic parity or parity of true positive rates), via FairGAN, DECAF, or TabFairGAN.

  • Fraud detection and class-imbalanced tabular problems. Generating synthetic samples to alleviate class imbalance, and using tabular generators such as CTGAN, CTAB-GAN, and TabDDPM for mixed-type tabular data.

Industry relevance. The paper notes that synthetic data are already pervasive in LLM training pipelines and are used across scientific, industrial, and policy domains, from augmenting medical imaging datasets to improve diagnostic model performance to generating samples for fraud detection. Because these deployments increasingly inform consequential decisions, understanding when synthetic data support valid inference — and where synthesis error propagates — is directly relevant to model developers, data governance teams, and analysts working under privacy regulation.

Future Directions

  • Theory for how synthesis error and misspecification propagate. The paper calls for understanding how errors from misspecified generative models flow through inferential and predictive workflows, so that robustness to misspecification can be characterized rather than assumed.

  • Statistically principled frameworks for using synthetic data. The paper proposes to organize and compare paradigms for principled synthetic data use and to delineate the assumptions under which each can provide plausible guarantees. It notes that new statistical theory and methodology are needed to leverage high-fidelity generations for reliable scientific discovery.

  • Correct handling of uncertainty. Attenuated uncertainty is listed among the common pitfalls that arise when synthetic data are treated as surrogates for real observations. How to propagate synthesis uncertainty correctly across the different access patterns remains an open methodological question.

  • Calibration and weighting protocols for combining real and synthetic data. Since naive pooling leads to invalid inference and bias, and since the augmentation access pattern uses O ∪ S "with potential calibration or weighting steps," determining appropriate calibration and weighting is an open practical problem.

  • Uses of synthetic data in LLM pre-training. The paper states directly that while utility in post-training stages such as instruction-tuning and alignment is fairly established, uses and effects in pre-training are poorly understood, and that uncontrolled recursive training causes model collapse.

  • Practical recommendations and cautions. The paper's stated conclusion includes practical recommendations and cautions to guide both method developers and applied researchers; the specific contents are not included in the truncated text.

Target Audience

Statisticians and biostatisticians working on inference, missing data, privacy, or fairness; machine learning researchers who build or evaluate generative models intended for downstream use; applied researchers in healthcare, biomedical multi-site studies, and the social sciences who use synthetic data for augmentation, imputation, or privacy-preserving release; and practitioners and data governance teams in industry who deploy generative models in LLM pipelines, tabular synthesis, fraud detection, or fairness-sensitive decision systems. Readers without a statistical background may find the notation in Section 2 and the differential privacy formulation challenging, but the five-motivation framework and the model-class comparison table are accessible on their own.

Authors’ abstract

The emergence of generative AI models has dramatically expanded the availability and use of synthetic data across scientific, industrial, and policy domains. While these developments open new possibilities for data analysis, they also raise fundamental statistical questions about when synthetic data can be used in a valid, reliable, and principled manner. This paper reviews the current landscape of synthetic data generation and use from a statistical perspective, with the goal of clarifying the assumptions under which synthetic data can meaningfully support downstream discovery, inference, and prediction. We survey major classes of modern generative models, their intended use cases, and the benefits they offer, while also highlighting their limitations and characteristic failure modes. We additionally examine common pitfalls that arise when synthetic data are treated as surrogates for real observations, including biases from model misspecification, attenuated uncertainty, and difficulties in generalization. Building on these insights, we discuss emerging frameworks for the principled use of synthetic data. We conclude with practical recommendations, open problems, and cautions intended to guide both method developers and applied researchers.

Read the original paper