Skip to content
AI.info

Research

Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport

Overview Research area: Robust machine learning and out-of-distribution (OOD) data synthesis, combining entropic (Sinkhorn) optimal transport, distributionally robust optimization (DRO), and latent-sp

Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport
arXiv
2609.31470
Published
2026-09-25
Authors
Haixiang Sun, Andrew L. Liu

AI summary

Overview

Research area: Robust machine learning and out-of-distribution (OOD) data synthesis, combining entropic (Sinkhorn) optimal transport, distributionally robust optimization (DRO), and latent-space generative modeling. The paper is posted under stat.ML.

Technical level: Advanced. The paper is built on Wasserstein/Sinkhorn DRO duality, Gibbs kernels, total-variation convergence bounds, and finite-sample generalization arguments, and it presumes familiarity with these tools.

Scope: The paper proposes Sinkhorn Boundary Outlier Generation (SBOG), a latent-space framework that generates synthetic outliers by scoring candidate points with a transport-based "Sinkhorn Outlier Energy" (SOE) and constraining them to remain semantically anchored, and it evaluates the resulting outliers through downstream time-series anomaly detection benchmarks and an image outlier-synthesis setting.

What This Paper Is About

Outliers are the cases that decide whether a model fails in deployment, but they are rare, varied, and poorly covered by finite datasets, so simple resampling or perturbation cannot produce them. Existing synthesis methods select candidates using heuristics such as sparse nearest-neighbor neighborhoods, low-likelihood latent regions, or classifier-boundary crossings, which the authors argue are unstable and tied to specific modalities or architectures. This paper reframes outlier synthesis as constrained latent boundary generation: candidates must be weakly supported by the in-distribution reference measure under a Sinkhorn transport geometry, while remaining semantically anchored and generatively decodable.

Key Contributions

  1. A new problem formulation. The authors formulate structured outlier generation as constrained latent boundary generation, where candidates are required to be weakly supported by the in-distribution reference measure while remaining semantically anchored and generatively valid. This explicitly reframes outlier synthesis away from low-density or sparse-neighborhood search.

  2. The SBOG algorithm and Sinkhorn Outlier Energy (SOE). SBOG uses SOE to select boundary anchors, accept weakly supported candidates under the Sinkhorn transport geometry, enforce a semantic feasibility constraint, and decode synthetic outliers. The authors show that in the empirical Euclidean case SOE reduces to a kernel-support score, but stress that SBOG uses it as a transport-based boundary criterion rather than as a standalone density estimator.

  3. A DRO-dual justification plus theoretical guarantees. The paper shows that SOE appears as the Gibbs-normalization term in the Sinkhorn-DRO dual, linking SBOG's thresholding and energy-maximizing selection to boundary decisions in a robust-optimization view. It supplements this with a finite-sample stability result for SOE thresholding (Theorem 3.4) and a total-variation consistency result relative to the population-SOE sampler (Theorem 3.5).

  4. Cross-modality empirical evaluation. Across time-series and image benchmarks, the authors report that SBOG generates informative, semantically controlled outliers and improves downstream robustness evaluation (detector-side discrimination), using downstream utility rather than a fidelity metric because synthetic outliers have no canonical target distribution.

Main Findings

  • SOE recovers a kernel-support form only in a restricted case. Lemma 3.2 shows that for the empirical reference measure ν̂_n = (1/n) Σ δ_{z_i}, SOE equals a minimum over the simplex of Σ w_i d(z, z_i) + ε Σ w_i log(n w_i), with softmax weights as the minimizer. The authors state that only in the special case of uniform empirical support with squared Euclidean cost does SOE reduce to a monotone transform of an unnormalized RBF-KDE score; they report that beyond this setting SOE is not equivalent to KDE and that their experiments show it yields stronger outlier anchors than KDE-based alternatives.

  • SOE is the support-normalization term of the Sinkhorn-DRO dual. Proposition 3.3 decomposes the one-dimensional Sinkhorn-DRO dual objective into a radius term plus a log-moment term plus λ · E^ν_ε(z), so the same outlier energy that SBOG thresholds also appears in the robust-optimization objective. The authors frame this as justification for the acquisition rule, not as claiming SBOG solves a task-specific DRO problem.

  • Finite-sample thresholding is stable under explicit conditions. Theorem 3.4 bounds the deviation of the empirical SOE from the population SOE by (2ε / K̄) · τ_n(δ) with τ_n(δ) = sqrt(log(2|V|/δ) / (2n)), and shows the threshold decision is preserved with probability at least 1 − δ when n is at least the maximum of (2/K̄²)·log(2|V|/δ) and (8ε²/(K̄²Δ_E²))·log(2|V|/δ). The sufficient sample size scales as O((ε²/(K̄²Δ_E²)) log(|V|/δ)); the authors note the bound can become loose if ε is too small or if the candidate pool sits in extremely weakly supported regions. They state that the sample complexity of neighborhood-based methods is analyzed in Appendix C, which is not included in the provided content.

  • Sampling is consistent with the population-SOE criterion. Theorem 3.5 gives a total-variation bound d_TV(Q_n^SBOG, Q*_{n,SOE}) ≤ min{1, 2{a_n + b_n(e_n)}/p_0} =: ζ_n, and the bound tends to zero in probability under stated conditions. The authors are explicit that this is consistency with the population SOE criterion, not alignment with an independently defined target boundary.

  • SBOG beats other CARLA negative-construction rules on time series. Using CARLA as a controlled testbed and changing only the negative-sample construction rule, average AU-PR and F1 across SMD, MSL, SMAP, SWaT, and WADI are: CARLA 0.267 AU-PR / 0.362 F1; CARLA-RandomNeg 0.235 / 0.340; CARLA-k-NNNeg 0.291 / 0.390; CARLA-KDENeg 0.280 / 0.393; CARLA-SBOG 0.312 / 0.432. The authors report CARLA-SBOG as the best average among all CARLA variants.

  • General time-series baselines are also compared. The upper part of Table 1 reports AnomalyTransformer (average 0.171 AU-PR / 0.268 F1), DCdetector (0.150 / 0.232), TranAD, COUTA, TS2Vec, ScatterAD, THOC (0.227 / 0.340), and Sub-Adjacent Transformer (0.299 / 0.410), evaluated with point-wise F1 and AU-PR without Point Adjustment, which the authors describe as the stricter setting.

  • Ablations isolate the sampler's mechanisms. Ablating high-energy anchor selection, energy-threshold acceptance, and argmax candidate selection, and comparing against k-NN and KDE support-based variants on the same trained encoder, panels (a) and (b) show SBOG keeps high anchor fidelity while inducing a stronger distributional shift as measured by MMD², whereas random sampling and removing anchor selection reduce fidelity. Panel (c) reports that SBOG attains a better margin–diversity trade-off than the ablated and support-score variants.

  • The energy threshold τ_E has a clear operating range. Table 2 reports, for τ_E values 0.50, 0.70, 0.90, 0.95, 0.98, and 0.99, average AU-PR of 0.279, 0.295, 0.310, 0.312, 0.305, and 0.277, average F1 of 0.387, 0.409, 0.430, 0.432, 0.422, and 0.384, and Gap values of 0.225, 0.224, 0.217, 0.126, 0.109, and 0.079, with Fill values of 1.000, 1.000, 1.000, 1.000, 0.999, and 0.945. The best values appear at τ_E = 0.95. Gap is defined as the average margin E(neg) − τ_E; the definition of Fill is cut off in the provided content.

  • The robust alignment module is not used for time series. The authors state that the Sinkhorn-DRO robust alignment module is omitted in the time-series experiments because those benchmarks contain no semantic class anchors or prototype labels defining logits T(c)ᵀz. Robust alignment reduces to the standard alignment loss when r_DRO = 0 and ν_x = δ_x (Remark 3.6).

  • Image results are claimed but not detailed in the provided content. The abstract states that experiments cover image outlier synthesis alongside time series anomaly generation, but the truncated content does not report the image benchmark numbers, dataset names, or metrics.

Methodology in Plain English

The authors work in a learned latent representation space rather than directly in pixel or sensor space. An encoder maps inputs to normalized embeddings, and for each embedding the method computes a quantity called Sinkhorn Outlier Energy: a smoothed measure of how much in-distribution mass lies nearby, weighted by a Gibbs kernel exp(−d/ε) over a ground cost d. Low energy means the candidate is expensive to support from the in-distribution data; high energy means it is well supported.

The SBOG sampler then proceeds in four steps. First, it computes this energy for every in-distribution embedding and keeps the top-A values as boundary anchors. Second, around each selected anchor it draws M isotropic Gaussian perturbed proposals and projects them back into the latent space if needed. Third, a proposal is accepted only if it crosses an energy contour Ê(v) ≥ τ_E and satisfies a semantic feasibility constraint cos(v, t_{y_i}) ≥ τ_sem against the class anchor of the source embedding — the semantic floor prevents the sample from drifting into arbitrary sparse regions. Fourth, among feasible proposals it takes the single candidate with the largest energy and decodes it back to input space through a fixed conditional generator, repeating until N_ood outliers are produced.

To connect this rule to robust optimization, the authors show that the same energy term appears inside the one-dimensional dual of a Sinkhorn distributionally robust optimization problem, which means the boundary decisions correspond to deviations that are expensive to support under the transport uncertainty geometry. They also introduce an optional robust alignment loss that replaces pointwise similarity logits with logits computed under a local Sinkhorn neighborhood, so that the learned representation geometry is better calibrated for boundary generation.

Evaluation follows the outlier-synthesis literature's convention of measuring downstream utility rather than fidelity: better synthesized outliers should carve a clearer boundary between inliers and outliers and thus improve detector discrimination.

Why This Matters

For research, the paper offers a principled alternative to the low-density heuristic for outlier synthesis, ties a sampling rule to a DRO dual, and provides stability and consistency guarantees for the thresholding decision. It also gives a modality-agnostic construction, since the method only needs a latent reference measure and a ground cost, which contrasts with pipelines tailored to time series or images alone. A practical caveat the paper itself raises is that low support is not sufficient: sparse points may be sparse because they are informative boundary cases, or because they have simply drifted off the semantic manifold, and SBOG's semantic constraint is designed to distinguish the two.

Real-world applications implied by the paper's framing and benchmarks:

  • Industrial and infrastructure monitoring. SWaT and WADI are among the time-series benchmarks used, representing sensor streams where rare failures determine whether monitoring systems detect problems in time.
  • Spacecraft and remote-sensing telemetry. MSL and SMAP are among the benchmarks used, and the paper motivates the work by noting that future distributions may differ significantly from historical training data.
  • Pre-deployment stress testing of machine-learning systems. The paper describes its output as a foundation for stress scenario generation beyond empirical support, which can reveal failure modes before deployment.
  • Image-based robustness evaluation. The abstract claims image outlier synthesis results, so the same pipeline is intended to support vision settings, though the specific image benchmarks are not reported in the provided content.

Industry relevance centers on the gap between finite training data and rare deployment conditions: the framework provides a way to manufacture informative, class-consistent rare cases on demand rather than waiting to observe them. The code is released publicly at https://github.com/hSun08/SBOG.

Future Directions

  • Task-aware variants. The authors state that when a downstream loss is available, the same Sinkhorn-DRO dual can be used to derive task-aware variants that augment the acquisition with a task

Authors’ abstract

Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier boundary crossings, which can be heuristic, unstable, and tied to specific modalities or architectures. We therefore propose Sinkhorn Boundary Outlier Generation (SBOG), a structured framework for latent-space outlier generation that couples Sinkhorn optimal transport geometry with distributionally robust boundary modeling. The resulting Sinkhorn-induced support cost guides the sampler toward weakly supported boundary regions, while semantic constraints prevent uncontrolled drift from the intended context, yielding controlled deviations from the in-distribution reference measure rather than arbitrary sparse-region samples. Experiments on time series anomaly generation and image outlier synthesis show that our framework produces informative, semantically controlled outliers and improves downstream robustness evaluation across modalities, providing a foundation for stress scenario generation beyond empirical support.

Read the original paper