Skip to content
AI.info

Research

Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction

Overview Research area: Enterprise AI — synthetic database generation and tool-calling agent training data. Technical level: Intermediate (accessible with basic familiarity with LLMs, agents, and data

Synthesis Through Simulation: Generating Coherent Enterprise Data via Scalable Agent-System Interaction
arXiv
2610.10549
Published
2026-09-24
Authors
Yipeng Li, Ashutosh Hathidara, Jane Lo, Harshavardhan Abichandani, Gunraj Singh, Atin Ghosh

AI summary

Overview

Research area: Enterprise AI — synthetic database generation and tool-calling agent training data. Technical level: Intermediate (accessible with basic familiarity with LLMs, agents, and databases; the paper is formal in places and the full content includes metric derivations in appendices). Scope: The paper proposes a paradigm for generating structurally valid, distribution-faithful enterprise database snapshots by having an LLM agent populate a simulated enterprise system through its policy-enforcing APIs, and evaluates it across ten domains.

What This Paper Is About

Training and evaluating tool-calling agents requires realistic enterprise database snapshots, but real production data and database schemas are usually blocked by privacy, compliance, and governance restrictions. Existing synthetic tabular generators reproduce statistical distributions but cannot enforce relational, policy-defined business rules and need large seed corpora, while procedure-based generators enforce rules but require hand-authoring per domain and do not model realistic value distributions. The paper's goal is to generate enterprise data that is both structurally valid and distributionally faithful, at scale, without access to database schemas or per-domain configuration.

Key Contributions

  1. A formal problem decomposition of enterprise data synthesis into three independent axes — structural validity, distributional fidelity, and scalability — showing that generation at the API layer guarantees structural validity by construction (Proposition 1).
  2. Synthesis Through Simulation (STS), a schema-free synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs inside simulated enterprise environments, decoupling validity enforcement from distribution modeling.
  3. The Generalist Populator (GP), a zero-authoring, domain-agnostic agent that requires only the raw tool registry (function signatures and descriptions) — no database schema, no hand-authored intent library, no sampler, and no per-environment configuration.
  4. An open-source release of the full framework, all ten environments, and the generated datasets at https://github.com/SAP/synthesis-through-simulation.

Main Findings

  • Structural validity is guaranteed, not learned. GP achieves VPR = 1.0 across all ten environments, matching Proposition 1. In contrast, offline statistical methods reach at most 0.923 VPR because they bypass the API enforcement layer entirely.
  • High distributional fidelity without schema access. GP reaches 0.88 average marginal fidelity and 100% constraint satisfaction across all ten environments without DB schemas. Marginal fidelity exceeds 0.93 in five environments: E-commerce (0.983), Obfuscated Airline (0.977), Airline (0.973), IT Mgmt (0.966), Forum (0.945), and HR (0.939).
  • Two environments show low cross-table fidelity from task-ratio imbalance. Forum CDF is 0.318 and Banking CDF is 0.421, attributed to the planner assigning roughly equal weight to entity-creation and follow-on operation intents from LLM world knowledge. In Forum, thread-creation tasks occur as often as post-reply tasks, yielding mostly 1–2 posts per thread while the reference expects a median of 3 (tail to 10); in Banking, account creation is sampled at a similar rate to deposits and withdrawals while the reference expects 5–30 transactions accumulated per account.
  • Semantic independence. On an obfuscated airline variant with opaque identifiers, GP matches the semantic baseline on all three metrics; the two inferred dependency graphs share 11 nodes and differ by only 2 edges (GED = 2, GED_norm = 0.154).
  • Schema privilege does not buy distributional quality. Against EnvScaler, which receives the full schema, live state, and business rules, GP wins on all distributional metrics on airline while EnvScaler fails entirely with 82% zero-write trajectories. On cloud, CDF favors EnvScaler (0.649 vs. 0.610). On HR, the difference is minimal (MDF tied; CDF 0.505 vs. 0.509).
  • Statistical synthesizers are often inapplicable. Seven of ten environments have no pre-populated seed data, so CTGAN, TVAE, SDV HMA, and REaLTabFormer could only be compared on three environments (airline with 8 seed rows, retail with 10, banking with 20). On airline GP also leads on distributional fidelity, but falls slightly below on retail and banking.
  • Exploration budget of K = 3 is the practical sweet spot. PDF rises sharply from K = 0 (0.813) to K = 3 (0.900) then flattens; MDF stays largely flat (0.95–0.98, peak at K = 3); CDF stays flat through K = 10 then dips at larger budgets as over-diversified exploration shifts child-count distributions away from the reference prior.
  • The Plan phase is what steers distributional fidelity. Disabling Plan on airline (K = 3, N = 50) drops all metrics (MDF −0.065, PDF −0.062, CDF −0.071) despite the agent writing far more per trajectory (16.7 bookings vs. 0.61 with planning), because it bulk-iterates over all users and the round-trip proportion collapses from 52.5% to 7.9% against a 55% reference prior.
  • Heuristic extraction is load-bearing across workflow types. Without it, banking WO drops 33% and PDF drops 0.175 with 69% of trajectories collapsing to a stereotyped three-step sequence; HR zero-write failure rises from 0.07 to 0.12; airline Calls/WO nearly doubles (+96%) without a "list existing reservations" shortcut.
  • STS snapshots improve downstream agent training. Fine-tuning Qwen2.5-7B on rollouts from an STS-populated airline environment beats both the base model and SFT on the pre-populated seed snapshot across all five TED metrics on the τ²-bench airline benchmark. Pass@k goes from 0.286 (base) to 0.500 (SFT with STS) vs. 0.375 (SFT with Seed); PPT avg_max goes from 0.165 to 0.326 vs. 0.249. SFT (Seed) fails to surpass the base model on AUC (0.499 vs. 0.504) and Mean Progress (0.309 vs. 0.309), suggesting state diversity drives the gain.
  • Latency tradeoff is real but arguably not comparable. GP's per-trajectory latency is 41–118 seconds, lower throughput than offline synthesizers per row, but offline synthesizers require seed data that is often absent and using production PII to train them creates a compliance conflict.

Methodology in Plain English

The researchers build ten mock enterprise systems (human resources, airline reservations, IT service management, banking, e-commerce, retail, university enrollment, pharmacy dispensing, online forums, and cloud computing infrastructure), each backed by SQLite and exposing only a typed set of tools plus a validation suite of programmatic business-rule checks. The database schema is deliberately hidden, mirroring the real governance boundary where APIs are sanctioned but schemas are restricted.

An LLM agent, the Generalist Populator, populates each system by calling those tools. Because every write goes through the same layer that defines validity, any accepted trajectory produces a coherent snapshot automatically — this is the validity-by-construction argument. The agent runs as an outer loop of N trajectories, each a concrete task, and keeps a persistent SQLite memory called the ExplorationManifest.

Each trajectory has three phases. In Explore, which runs only for the first K trajectories, all write tools are blocked so the agent builds a mental model purely from read calls and tool signatures: it classifies tools as read or write, identifies entity types and creation order, and estimates how frequently each operation should appear. In Plan, a single structured LLM call produces a concrete task description, using recent task history to avoid repetition and a coverage check to flag unused write operations; it anchors attribute targets to reference distributions when supplied and otherwise relies on the model's world knowledge. In Execute, the agent runs a full tool-calling loop with unrestricted tool access, discovers real identifiers from the live database, calls write tools in dependency order, and adapts from structured rejection messages when a precondition fails. After each trajectory, a best-effort step reviews the call sequence and persists any late-discovered shortcut as a reusable named heuristic.

Evaluation compares generated distributions against reference distributions grounded in public sources (industry reports, government statistics) using three Jensen-Shannon-divergence-based metrics at increasing granularity: marginal per-column fidelity (MDF), pairwise within-table joint fidelity (PDF), and cross-table child-count fidelity (CDF). Numeric columns use the complement of the Kolmogorov–Smirnov statistic. All metrics report 95% bootstrap confidence intervals with B = 1000.

Why This Matters

The paper reframes enterprise data synthesis: instead of asking a model to imitate rows, it asks an agent to operate a system, which turns a hard constraint-satisfaction problem into a normal agent interaction problem. This lets structural validity and distributional realism be attacked separately, and it removes the two practical blockers of schema disclosure and seed data availability.

Real-world applications:

  • Tool-calling agent training and evaluation where production databases cannot be exported, providing safe starting snapshots for tasks like policy-constrained retail or airline workflows.
  • Enterprise software testing and QA, supplying coherent pre-populated multi-table states without touching customer data.
  • Developer onboarding and demo environments, generating plausible business data for domains such as HR payroll, IT service management, or banking transaction histories.
  • Pre-deployment benchmarking of agents against mock systems that enforce the same API-level business rules as production, before any real-system access is granted.

Industry relevance is direct: the work is done at SAP Labs, the environments are enterprise systems, and the open-sourced framework targets exactly the compliance-constrained settings that enterprise vendors face. The result that a schema-free agent matches or exceeds a fully schema-privileged baseline on distributional quality is the strongest practical claim for vendors who cannot share schemas.

Future Directions

  • Intent-weight calibration. Closing the residual CDF gap in Forum (0.318) and Banking (0.421) by inferring realistic operation priors from seed database statistics or domain-context prompting, rather than relying on unanchored LLM world knowledge.
  • Cross-environment heuristic transfer. The paper names transferring learned heuristics between environments via tool-signature similarity as a natural extension.
  • Handling model-specific bias and enforcement completeness. GP's fidelity is sensitive to model quality, and the two-tier distribution mechanism introduces model-specific biases for unanchored fields; the validity guarantee is also conditional on the API layer being complete with respect to business rules, which could demand significant engineering effort in a real deployment.
  • Extending beyond the current scope. Relational, API-mediated databases with explicit validation logic are covered; unstructured content and probabilistic enforcement are explicitly out of scope, leaving them open.

Target Audience

Researchers and engineers working on synthetic data generation, tool-calling agents, and enterprise AI infrastructure will get the most from this paper. It is also relevant to practitioners who need to produce safe, realistic database snapshots for agent training, evaluation, or testing under data-governance restrictions, and to readers interested in how environment design can make hard correctness properties free by construction. Readers looking for statistical tabular synthesis techniques or for procedural constraint-encoding approaches will find it useful as a comparison point, since the paper positions both against STS in Table 1.

Authors’ abstract

Tool-calling agents have become central to enterprise AI, yet training and evaluating them at scale remains severely constrained due to business and legal restrictions on enterprise systems, data, and database schemas. Tabular data synthesis offers a natural alternative, but its effectiveness is fundamentally limited by structural validity and schema availability, while procedure-based approaches yield the opposite weakness, typically lacking distributional fidelity without per-domain authoring. We introduce **Synthesis Through Simulation** (STS), a **schema--free** data synthesis paradigm in which an LLM agent generates data by executing operations against policy-enforcing APIs within simulated enterprise environments. Because data is generated through the same environment that defines what is valid, STS guarantees structural validity by construction while decoupling validity enforcement from distribution modeling, allowing each to be addressed independently. The **Generalist Populator** (GP), STS's domain-agnostic agent, addresses the remaining challenges of distributional fidelity and synthesis scalability: GP achieves **0.88** average marginal fidelity and **100\% constraint satisfaction** across all ten environments *without access to DB schemas*, while statistical synthesizers are inapplicable to seven due to necessary seed data requirements, and schema-privileged agents fail 82\% of trajectories on airline environment's tightly coupled workflows due to brittle task composition. We open-source the full framework, all ten environments, and generated datasets at https://github.com/SAP/synthesis-through-simulation.

Read the original paper