Skip to content
AI.info

Research

Privacy-Aware Time Series Synthesis via Public Knowledge Distillation

Privacy-Aware Time Series Synthesis via Public Knowledge Distillation Overview Research area: Privacy-preserving machine learning, specifically differentially private synthetic time series generation

Privacy-Aware Time Series Synthesis via Public Knowledge Distillation
arXiv
2511.00700
Published
2025-11-01
Authors
Penghang Liu, Haibei Zhu, Eleonora Kreacic, Svitlana Vyetrenko

AI summary

Privacy-Aware Time Series Synthesis via Public Knowledge Distillation

Overview

  • Research area: Privacy-preserving machine learning, specifically differentially private synthetic time series generation with conditional diffusion models.
  • Technical level: Intermediate. The paper assumes familiarity with differential privacy (DP), DP-SGD, and diffusion models, though the core idea of using public "context" data is explained accessibly.
  • Scope: The paper proposes Pub2Priv, a conditional diffusion framework that generates DP-protected time series by conditioning on heterogeneous public metadata, and evaluates it on three domains (investment portfolios, electricity usage, semiconductor trading).

What This Paper Is About

Sharing sensitive time series — investment portfolios, household electricity usage, patient records — is restricted because individual records can be re-identified. The standard fix, differential privacy during training, forces a hard trade-off: more noise means more privacy but less useful synthetic data. This paper argues that this trade-off can be improved by exploiting an overlooked resource: public, non-sensitive contextual signals that are correlated with the private data (weather and electricity prices for consumption; the S&P 500 and Dow Jones for portfolios; the Philadelphia semiconductor index for chip imports). The goal is to generate private time series conditioned on those public signals without spending any additional privacy budget.

Key Contributions

  1. A new problem formulation for privacy-aware time series generation that treats publicly available, heterogeneous metadata as a conditioning signal rather than as an extra dataset — a setting the authors state no prior work has explored for DP time series generation.
  2. Pub2Priv, a conditional diffusion framework in which a pre-trained transformer encodes multi-dimensional public metadata into temporal and feature embeddings that condition the denoiser. Because the transformer never touches private data, it adds no privacy cost.
  3. A practical privacy metric based on the identifiability of synthetic data, proposed as a more interpretable way to assess and configure privacy guarantees for generation models.
  4. An empirical evaluation across three domains — finance, energy, and commodity trading — against DP-GAN, PATE-GAN, GEM, AIM, and Private-GSD, including ablations without metadata and without the knowledge transformer.

Main Findings

  • Better distributional fidelity on returns and autocorrelation: Under a privacy budget of ε = 1, δ = 1×10⁻⁵, Pub2Priv reaches KS_R of 0.28 ± 0.02 (Portfolio), 0.26 ± 0.01 (Electricity), and 0.32 ± 0.03 (Comtrade), and KS_AR of 0.17 ± 0.05, 0.17 ± 0.02, and 0.77 ± 0.01 respectively. The paper states Pub2Priv achieves the best KS_R on the Portfolio dataset and better or comparable KS_R on the other two, and that it consistently outperforms all baselines on KS_AR across datasets.
  • Note on one metric: On the Comtrade dataset, the reported KS_R values for GEM (0.19 ± 0.00), AIM (0.20 ± 0.00), and Private-GSD (0.22 ± 0.00) are lower than Pub2Priv's 0.32 ± 0.03, even though the table marks Pub2Priv in bold as best "among Pub2Priv and the baselines."
  • Best preservation of private–public correlations: Pub2Priv yields the smallest |Corr_meta| on every dataset — 0.18 ± 0.01 (Portfolio), 0.08 ± 0.02 (Electricity), 0.44 ± 0.05 (Comtrade). The paper notes this metric is diagnostic and is not part of either model's training objective.
  • Best downstream utility on average: The discriminative TSTR score improves by 0.118 and the predictive score by 0.013 compared to baselines. Pub2Priv's TSTR (discri.) is 0.100 ± 0.08 (Portfolio), 0.119 ± 0.066 (Electricity), and 0.113 ± 0.032 (Comtrade); TSTR (predic.) is 0.003 ± 0.000, 0.007 ± 0.007, and 0.013 ± 0.001.
  • One exception to the TSTR claim: DP-GAN achieves the best discriminative TSTR on the Electricity dataset (0.111 ± 0.145), but the paper attributes this to instability, noting its standard deviation is larger than its mean.
  • Public metadata is the main driver: Ablations removing the conditioning metadata (Pub2Priv w/o c) degrade sharply — for example |Corr_meta| rises from 0.18 to 0.53 on Portfolio and KS_AR from 0.17 to 0.42 — while removing only the transformer (Pub2Priv w/o θ_T) hurts less, indicating the metadata signal matters more than the specific encoder.
  • Marginal-statistics methods fail at individual-level fidelity: GEM, AIM, and Private-GSD perform poorly on TSTR because they are designed for accurate answers to aggregate queries rather than realistic individual trajectories such as plausible investment portfolios.
  • Qualitative coverage: In t-SNE visualizations (ε = 1, δ = 1×10⁻⁵), DP-GAN, PATE-GAN, and AIM deviate substantially from the real distribution, while GEM and Private-GSD cluster around a few outliers; Pub2Priv samples more closely match the real distribution. t-SNE plots for the Comtrade dataset are omitted because its relatively small size produces sparse visualizations.
  • Trend across privacy budgets: Across privacy budgets, Pub2Priv consistently yields better TSTR scores than baselines (Figure 3, Portfolio dataset). All methods struggle at very tight budgets, and Pub2Priv increasingly captures structure tied to public signals as ε approaches 1.

Methodology in Plain English

Pub2Priv separates the problem into "public knowledge" and "private generation."

First, a pre-trained transformer reads only public data. It uses a two-dimensional self-attention design that alternates between attending along time (within a single signal) and along features (across signals at a single time step), which models long-range temporal structure and cross-signal correlations efficiently. Sinusoidal temporal position encodings and learned feature embeddings are added before encoding, and attention masks handle variable-length sequences and missing entries. The transformer is pre-trained with a masked reconstruction objective on publicly available stock prices and weather from Yahoo Finance and NCEI, minimizing a per-feature normalized mean-squared error over masked entries. After pre-training it is frozen, and its output — one embedding vector per time step — becomes the conditioning signal.

Second, a diffusion model generates the private time series. During the forward process, Gaussian noise is progressively added to real samples; during the reverse process, the denoiser predicts the noise while being conditioned on the frozen public embeddings at every diffusion step. Conditioning is applied without adding any noise.

Third, privacy is enforced only where private data is involved. Training of the denoiser uses DP-SGD: per-sample gradients are clipped to an ℓ₂ threshold C (the authors search C ∈ {0.1, 0.5, 1.0, 1.5, 2.0} and pick the value with lowest validation loss), Gaussian noise 𝒩(0, σ²C²I) is added to the averaged gradient, and parameters are updated with the noisy gradient. Cumulative privacy loss is tracked with a moments' accountant or Rényi Differential Privacy. Because the transformer is pre-trained without access to private data, the paper argues it incurs no privacy loss, so (ε, δ)-DP holds for the overall model.

Evaluation compares Pub2Priv against DP-GAN, PATE-GAN, GEM, AIM, and Private-GSD under the same budget (ε = 1, δ = 1×10⁻⁵), with each experiment repeated 10 times and mean ± standard deviation reported. Utility is measured by KS statistics on return and autocorrelation distributions, the metadata-correlation gap, and Train-on-Synthetic-Test-on-Real discriminative and predictive scores. Privacy is assessed by the proposed identifiability metric, computed from nearest-neighbor distances between synthetic samples and real data.

Why This Matters

The paper reframes a limiting assumption in privacy-preserving generation: that useful auxiliary data must come from the same distribution as the private data. In real deployments, homogeneous public data is scarce, but correlated public context — weather, prices, market indices — is abundant and free of privacy concerns. Showing that such context improves the privacy-utility trade-off at zero additional privacy cost gives practitioners a concrete lever that does not require weakening DP guarantees.

Real-world applications:

  • Financial data sharing: Generating realistic synthetic portfolio and trading sequences that preserve individual-level dynamics, allowing model development and stress-testing without exposing client positions.
  • Energy analytics: Sharing household electricity consumption patterns for grid planning and demand forecasting while using temperature and pricing data as public conditioning signals.
  • Trade and supply-chain analysis: Producing synthetic import/export series correlated with public sector indices (e.g., the SOX index) for research on semiconductor trade flows.
  • Regulated data collaboration: Enabling synthesis pipelines where the sensitive data cannot leave an institution but correlated public signals can, as a practical complement to anonymization approaches.

Industry relevance is direct: the work comes from JPMorgan AI Research, and the datasets and framing — portfolios, DP budgets, query-release baselines — reflect the constraints of regulated financial institutions that must publish or share data-like artifacts without exposing clients.

Future Directions

  • Extending beyond the three tested domains. Healthcare is used as a motivating example in the abstract but is not among the evaluated datasets, so the framework's behavior on clinical or patient-record time series remains untested.
  • Automatic discovery of public metadata. The paper cites prior work (Zhu et al., 2024) on using knowledge embedded in large language models to identify auxiliary datasets; automating which public signals to select and align with a given private dataset is an open step.
  • Generalizing metadata types. The formulation allows c to be categorical or continuous, but the implementation and experiments use time-varying continuous metadata only; how the framework handles categorical or irregularly sampled public signals is not reported.
  • Further validation of the identifiability metric. The new privacy metric is introduced as a practical, interpretable alternative, but the paper's comparison against established accounting-based guarantees and against other privacy mechanisms is limited.
  • Understanding where conditioning hurts. The Comtrade results, where marginal-statistics baselines post lower KS_R than Pub2Priv, and the Electricity case where DP-GAN wins the discriminative TSTR despite high variance, raise questions about when public conditioning helps and when simpler methods suffice.

Target Audience

Researchers and practitioners working on differentially private synthetic data, generative time series modeling, or privacy-preserving machine learning — particularly those in finance, energy, and other regulated settings who need to share or publish data-derived artifacts. It is also relevant to readers interested in conditional diffusion models and in how external, non-sensitive context can be incorporated into generative training without expanding a privacy budget. Readers without background in differential privacy will find the introduction and problem formulation accessible, but the privacy analysis and metric details are suited to a more technical audience.

Authors’ abstract

Sharing sensitive time series data in domains such as finance, healthcare, and energy consumption, such as patient records or investment accounts, is often restricted due to privacy concerns. Privacy-aware synthetic time series generation addresses this challenge by enforcing noise during training, inherently introducing a trade-off between privacy and utility. In many cases, sensitive sequences is correlated with publicly available, non-sensitive contextual metadata (e.g., household electricity consumption may be influenced by weather conditions and electricity prices). However, existing privacy-aware data generation methods often overlook this opportunity, resulting in suboptimal privacy-utility trade-offs. In this paper, we present Pub2Priv, a novel framework for generating private time series data by leveraging heterogeneous public knowledge. Our model employs a self-attention mechanism to encode public data into temporal and feature embeddings, which serve as conditional inputs for a diffusion model to generate synthetic private sequences. Additionally, we introduce a practical metric to assess privacy by evaluating the identifiability of the synthetic data. Experimental results show that Pub2Priv consistently outperforms state-of-the-art benchmarks in improving the privacy-utility trade-off across finance, energy, and commodity trading domains.

Read the original paper