Research
Developing synthetic microdata through machine learning for firm-level business surveys
Overview Research area: Privacy-preserving synthetic data generation for official statistics, applied to firm-level business surveys using machine learning (classification and regression tree synthesi

- arXiv
- 2512.05948
- Published
- 2025-12-05
- Authors
- Jorge Cisneros, Timothy Wojan, Matthew Williams, Jennifer Ozawa, Robert Chew, Kimberly Janda, Timothy Navarro, Michael Floyd, Christine Task, Damon Streat
AI summary
Overview
Research area: Privacy-preserving synthetic data generation for official statistics, applied to firm-level business surveys using machine learning (classification and regression tree synthesizers).
Technical level: Intermediate. The paper combines accessible conceptual explanation with econometric replication tables, but readers should be comfortable with concepts like principal component analysis, odds ratios, and regression coefficient confidence intervals.
Scope: The paper describes a machine learning framework for generating synthetic public-use microdata (PUMS) from confidential firm-level business surveys, and evaluates the fidelity of that synthetic data by replicating published econometric research.
What This Paper Is About
Public-use microdata from the US Census Bureau has long been available for individuals, but firm-level business survey microdata remains restricted because detailed industry and geographic information makes businesses easy to re-identify. The authors describe a machine learning approach that generates artificial firms preserving the statistical properties of the real data without containing any actual respondent records. Because the Annual Business Survey (ABS) synthetic results are still confidential at the time of submission, the paper demonstrates the approach on the 2007 Survey of Business Owners (SBO) PUMS, a comparable and unrestricted firm-level dataset.
Key Contributions
-
A CART-based synthesis framework for business microdata. The paper describes CenSyn, which consists of a Synthesizer (a sequence of decision trees) and an Evaluator suite, plus a parallel implementation using the R
synthpoppackage for sensitivity testing against software configuration. CenSyn handles privacy checks, consistency checks, survey weights, full or partial synthesis, and preservation of data sparsity. -
A fast distribution-similarity metric for iterative tuning. The k-marginal tool produces a score between 0 and 1000, where 0 implies no overlap between ground-truth and synthetic PUMS and 1000 implies perfectly matching density distributions. The paper reports that a score of at least 970 is sought empirically for a satisfactory synthetic PUMS.
-
Two synthetic versions of the 2007 SBO PUMS. Using the original 2007 PUMS (over 2.1 million firms and 200 features) as ground truth, the authors generate two synthetic PUMS of over 1 million synthetic firms each with 54 features — roughly half the number of firms of the original, for computational feasibility — one from CenSyn and one from synthpop.
-
An econometric replication test of verisimilitude. The authors replicate the analysis in Wang & Liu (2015), "Transnational activities of immigrant-owned firms and their performances in the USA," published in Small Business Economics, using both synthetic datasets, and separately report a failed replication of Mora & Dávila (2014) from the American Economic Review.
Main Findings
-
k-marginal scores were high and identical across frameworks: The synthpop-synthesized dataset scored 992, and the CenSyn-synthesized dataset also scored 992. Both had a 30% baseline score of 994.
-
State and revenue were the hardest features to synthesize: Based on other metrics, the state feature (
FIPST) and revenue (RECEIPTS_NOISY) were the worst performing features. The authors attribute the state difficulty to diversity of responses across the microdata, and the revenue difficulty to noise infused into the released 2007 PUMS — a complication they do not anticipate in the internal ABS data. -
Principal component structure was closely matched: The top five principal components explained 13.78%, 7.12%, 3.12%, 1.55%, and 1.35% of variance in the SBO 2007 PUMS, versus 13.78%, 7.09%, 3.11%, 1.53%, and 1.36% in the CenSyn PUMS. The components account for a quarter of total variance, with PC-1 contributing about 14% and PC-5 nearly 1%.
-
First moments were reproduced to the third decimal place for many variables: In Table 2, Exporting was 0.1150 in the 2007 SBO PUMS, 0.1161 in CenSyn, and 0.1156 in synthpop. International Branch was 0.0133, 0.0126, and 0.0131 respectively; International Outsourcing was 0.0164, 0.0164, and 0.0163.
-
Inter-feature distributions were largely preserved: The paper shows conditional distributions such as the age of the first owner of firms in Washington established between 2000 and 2005 that are jointly owned with a spouse but primarily by the husband, and the education level of the first owner of firms with startup capital between $5,000 and $10,000 in the "Professional, Scientific, & Technical Services" sector in Ohio.
-
Qualitative econometric conclusions matched, quantitative ones sometimes did not: Immigrant ownership was positively associated with all three transnational activities across all three datasets, and immigrant-owned family businesses were negatively associated with exporting in all three. However, the CenSyn immigrant estimates were statistically lower than the 2007 SBO estimates in the International Branch and Exporting equations, while the synthpop estimates were statistically lower in the International Outsourcing equation.
-
Large effects were attenuated in synthetic data: The effect of transnational business on sales per employee and payroll per employee was larger using the observational dataset than either synthetic dataset for the more prevalent Transnational feature (0.1272 share of firms), but not for the rarer immigrant-owned transnational business (0.0188). One estimate — the Immigrant coefficient in the Sales per Employee equation — was not statistically significant using CenSyn but was significant using the observational data or synthpop.
-
Industry fixed effects were well captured: Difficult-to-export industries such as Construction, Social Services, and Personal Services showed large negative estimates, while tradable industries such as Wholesale, Transportation, and Information had positive and statistically significant estimates, with Manufacturing as the excluded category. The synthetic datasets picked these up with generally similar magnitudes.
-
A replication of different research did not succeed: An attempt to replicate Mora & Dávila (2014) on the failure rate of firms in their first year of operation did not produce identical qualitative results using synthetic data. The authors attribute this partly to restricting analysis to firms in their first year, which reduced the sample size by more than a factor of ten.
Methodology in Plain English
The authors treat synthetic data generation as fitting a generative model to a dataset and then sampling a new dataset from that distribution. They use classification and regression trees (CART), which recursively split the data on features to form a tree of low-entropy, self-similar groups. Synthetic records are produced by following the same splitting rules learned in training, so the synthetic firms retain the underlying patterns and distributions of the real data without reproducing any actual record.
Two implementations are run in parallel to test sensitivity to software configuration: CenSyn (developed since 2018 by Knexus with the US Census Bureau) and the R synthpop library (created in 2016 as part of the UK SYLLS project). Evaluation proceeds in an iterative loop: an initial tuning round identifies the poorest performing features; results are shared with stakeholders and survey managers to hypothesize why certain features are harder to model; and a third phase explores solutions informed by that insight. The second and third phases repeat until a final product is achieved.
For validation, the authors compare distributions directly through histograms and bar plots, compare multivariate structure via pairwise PCA plots of the top five principal components, and then run the econometric specification from Wang & Liu (2015) on both the observed and synthetic datasets. The replication uses 26 independent variables, excludes non-employer businesses (since the exercise informs ABS, which covers only employer businesses), and omits three control variables that were not in the simulated feature set: whether the business is home-based, and whether the business sold to businesses or to consumers. Odds ratio confidence intervals are compared across datasets using the overlapping confidence intervals method of Schenker & Gentleman (2001).
Why This Matters
Impact on research: Synthetic data can give researchers a way to test code and specifications before submitting slow-moving requests to run on internal data, and to develop detailed pre-analysis plans. The authors argue this workflow — specification testing on synthetic data followed by a single hypothesis test on confidential data — can reduce researcher degrees of freedom and the probability of false discovery, an issue they describe as a principal contributing factor to the replicability crisis in applied economics research. They note the probability of false discovery is arguably even higher in innovation economics, where phenomena concentrate in the right tail of distributions and may require quantile regression to detect.
Real-world applications:
- Providing public-use firm microdata so that academic, government, and industry analysts can do exploratory research without obtaining Special Sworn Status, which requires in-depth background investigations with processing times of 3 to 4 months, and a requirement to have lived in the US for 36 of the past 60 months.
- Reducing reliance on the 34 Federal Statistical Research Data Centers for high-level analyses.
- Supporting reproducibility and transparency, as emphasized by the United Nations Economic Commission for Europe for testing code before submission to secure environments.
- Meeting the access-tier requirements of the US Foundations for Evidence-Based Policymaking Act of 2018, and the goals of the National Secure Data Service Demonstration Project authorized by the CHIPS and Science Act of 2022.
Industry relevance: The concern generalizes beyond government surveys. Re-identification risk rises as computing power and Big Data availability increase, and business records can be linked with public sources such as SEC filings, County Business Patterns, and National Establishment Time Series. The paper notes that the SBO 2007 PUMS is still requested today, indicating healthy demand for 18-year-old firm-level microdata and a potentially large latent demand for an update.
Future Directions
-
Producing the ABS PUMS. The long-term aim is a synthetic ABS PUMS via CART-based models. The authors were evaluating and fine-tuning results that remained restricted at the time of submission, with insights from the SBO work carried forward.
-
Handling small analytically important subsets. Since restricting analysis to firms in their first year reduced sample size by more than a factor of ten and was associated with the replication failure, the authors flag this as a topic of future research — especially relevant given strong research interest in R&D-performing microbusinesses in the ABS, which make up a small share of the survey.
-
Including additional features in the synthesized set. Whether the business is home-based and whether it sold to businesses or consumers were not simulated. The authors state that customer type questions are included in the ABS and will be important to include in the eventual synthetic public use file.
-
Refining the state and revenue features. The worst-performing features in the SBO exercise were the state code and revenue variables, and the authors expect the revenue difficulty to be an artifact of noise infused into the publicly released 2007 PUMS rather than a feature of the internal ABS data.
The provided paper content is truncated mid-sentence in the conclusion, so the full discussion of limitations and remaining future research is not available.
Target Audience
This paper is most useful for statisticians and data scientists at national statistical organizations and survey research firms who design public-use data products; economists and social scientists who rely on restricted firm-level microdata and want to know whether synthetic substitutes can support their analyses; privacy researchers evaluating synthetic data generation methods against alternatives such as cell suppression, subsampling, and differential privacy; and policy analysts interested in how Evidence Act and CHIPS and Science Act mandates for expanded data access may be satisfied technically.
Authors’ abstract
Public-use microdata samples (PUMS) from the United States (US) Census Bureau on individuals have been available for decades. However, large increases in computing power and the greater availability of Big Data have dramatically increased the probability of re-identifying anonymized data, potentially violating the pledge of confidentiality given to survey respondents. Data science tools can be used to produce synthetic data that preserve critical moments of the empirical data but do not contain the records of any existing individual respondent or business. Developing public-use firm data from surveys presents unique challenges different from demographic data, because there is a lack of anonymity and certain industries can be easily identified in each geographic area. This paper briefly describes a machine learning model used to construct a synthetic PUMS based on the Annual Business Survey (ABS) and discusses various quality metrics. Although the ABS PUMS is currently being refined and results are confidential, we present two synthetic PUMS developed for the 2007 Survey of Business Owners, similar to the ABS business data. Econometric replication of a high impact analysis published in Small Business Economics demonstrates the verisimilitude of the synthetic data to the true data and motivates discussion of possible ABS use cases.