Research
On the Information Processing of One-Dimensional Wasserstein Distances with Finite Samples
Overview Research area: Machine learning theory and optimal transport, with applied case studies in computational neuroscience and molecular biology. Technical level: Advanced. The paper is built on a

- arXiv
- 2511.12881
- Published
- 2025-11-17
- Authors
- Cheongjae Jang, Jonghyun Won, Soyeon Jun, Chun Kee Chung, Keehyoung Joo, Yung-Kyun Noh
AI summary
Overview
Research area: Machine learning theory and optimal transport, with applied case studies in computational neuroscience and molecular biology.
Technical level: Advanced. The paper is built on analytic derivations involving Poisson processes, Erlang distributions, binomial expectations, and closed-form one-dimensional optimal transport, so comfortable familiarity with probability theory is assumed.
Scope: The paper analyzes, both analytically and empirically, how the one-dimensional Wasserstein distance between finite samples encodes pointwise density (rate) differences and support differences.
What This Paper Is About
The Wasserstein distance is widely used because it measures differences between probability distributions through the transport cost between samples, which makes it sensitive to differences in support. What has been unclear is whether, and how, that same transport information can identify pointwise density differences when two distributions share nearly the same support but differ in their densities — and whether an analytic characterization exists in the finite-sample regime. The authors tackle this question in one dimension, where optimal transport has a closed form, and derive how expected sample transport distances respond to rate differences, support shifts, and both combined.
Key Contributions
-
An analytic characterization of rate encoding. Using Poisson processes with constant rates λ₁ and λ₂, the authors derive a closed-form expression for the expected distance between the k-th spikes of two processes (Proposition 3.1), showing it depends only on λ₁ and λ₂ and is symmetric in them.
-
Proof that the distance is minimized at equal rates. Under the constraint that the harmonic mean of the rates is held constant, the expected distance between the k-th spikes is minimized when λ₁ = λ₂, establishing that finite-sample Wasserstein distances reliably capture rate (density) differences.
-
A unified treatment of rate and support differences. By shifting the support of one process by Δt, the authors derive an expression (Equation 7 for the first spikes) in which rate-difference information and support-shift information are combined, with the balance between them moderated by the magnitude of Δt.
-
An extension sketch to time-varying rates and empirical validation. The paper provides a substitution-integral formulation for nonhomogeneous Poisson processes with time-varying rates μ(t) and ν(t), and confirms the theoretical predictions on synthetic data, salamander retinal ganglion cell spike trains, human hippocampal spike trains, and amino acid contact frequency data, connecting the findings to sliced Wasserstein distances.
Main Findings
-
Sample transport cost encodes both rate and support differences. In a synthetic prediction task using ten-dimensional transport-cost features computed over decile partitions of empirical measures, sample transport cost achieved R² scores of 81.5 ± 0.1 for log(r₁), 81.9 ± 0.2 for log(r₂), and 98.9 ± 0.0 for |Δt|. Directed Hausdorff features scored 43.7 ± 0.5, 43.9 ± 0.3, and 70.4 ± 0.3, and bin-wise JS divergence features scored 64.0 ± 0.4, 68.4 ± 0.3, and 70.3 ± 0.1 on the same targets.
-
The distance is minimized when rates are equal. Figure 2 shows E[W(μ̂_N, ν̂_N)] for λ₁, λ₂ ∈ [1, 5] and N = 20; along curves of constant harmonic mean of the rates (black dashed lines), the expected Wasserstein distance is smallest where λ₁ = λ₂. Both Wasserstein distance and individual sample distances such as |x₅₀ − y₅₀| capture and harmonize rate and support information, whereas Hausdorff fails to capture rate differences and JS divergence is overly sensitive for |Δt| < 1 and saturates for |Δt| ≥ 1 (values averaged over 1,000 trials).
-
A large-N approximation links rate differences to inverse-rate differences. For λ₁ < λ₂ and large N, the expected Wasserstein distance is approximated up to leading order as ((N+1)/2)(1/λ₁ − 1/λ₂), matching the Wasserstein distance between two uniform distributions U[0, (N+1)/λ₁] and U[0, (N+1)/λ₂]. The asymptotic per-spike result is lim_{k→∞} E[|x_k − y_k|/k] = 1/λ₁ − 1/λ₂ with vanishing variance.
-
As Δt grows, the support shift dominates. The expression in Equation (7) reduces to the pure rate term when Δt = 0, and to Δt + 1/λ₁ − 1/λ₂ as Δt → ∞.
-
Sample transport features improve neural stimulus classification. Adding SD1 or SD2 transport-cost features to inter-spike-interval (ISI) inputs improved test AUC across four 1D CNN architectures on salamander retinal ganglion cell data. Examples: FCN went from 0.945 ± 7e-04 to 0.951 ± 4e-04 (Retina-All) with SD1; XceptionTime went from 0.970 ± 1e-03 to 0.979 ± 6e-04 (Retina14) with SD1; ResNet improved from 0.937 ± 7e-04 to 0.948 ± 5e-04 (Retina-All) with SD2. Several of these gains are reported as statistically significant with p-values below 0.05.
-
Wasserstein embeddings reveal structure that spike count and Victor-Purpura distances miss. Isomap embeddings of human hippocampal spike trains (667-sec recording from four microelectrodes, segmented into 140-sec windows with 3.5-sec sliding interval, T = 151) showed a smoother trajectory and more consistent recall-related variation under the Wasserstein distance; the spike count difference and Victor-Purpura measures primarily distinguished the initial windows from the rest. Sharp jumps aligned with windows having extreme activity at their start or end (histograms beginning at 210.0 s, 213.5 s, 231.0 s, and 234.5 s).
-
Wasserstein embeddings of amino acids align better with hydrophobicity rankings. Using 12,508 proteins from the Protein Data Bank and histograms over 217 uniformly spaced bins, radial ordering in the Wasserstein embeddings correlated more strongly with the hydrophobicity rankings of Rose et al. (1985) than KL-based embeddings: Kendall's tau of 0.722 versus 0.582 for the top 10 most hydrophobic residues, and 0.807 versus 0.731 for all residues. CYS — identified as the most hydrophobic amino acid by Rose et al. (1985) — sits at the outermost edge in the Wasserstein embedding but closer to the center in the KL-based one. Large rate differences for CYS, ILE, and PHE were concentrated in the 8–12 Å range, indicating a propensity for long-range contacts.
Methodology in Plain English
The authors model the samples as event times generated by Poisson processes, the same mathematical object commonly used to describe neural spike timing. Because a Poisson process is fully described by its rate parameter, this setup lets the researchers isolate the effect of rate from every other factor. For a constant-rate process, the time of the k-th event follows an Erlang distribution, so the distance between the k-th events of two processes can be analyzed directly. The authors compute the expectation of the absolute distance between paired events and of the resulting empirical Wasserstein distance, which is simply the average of the matched-sample distances after sorting.
They then add a support shift Δt to one process and redo the calculation, which produces a formula in which the rate term is exponentially down-weighted, a Δt term appears, and a residual inverse-rate term remains. For time-varying rates, they substitute cumulative rate integrals for the time variable, converting the problem into an integral over two unit-rate Poisson variables.
Empirically, they validate these predictions on a synthetic dataset (comparing transport cost against Hausdorff distance and Jensen-Shannon divergence as features for a three-layer fully connected network), then apply the same transport-cost idea to neural spike train classification with four 1D CNN architectures, to Isomap embeddings of human hippocampal recordings, and to dissimilarity-based embeddings of amino acid contact distributions.
Why This Matters
The paper gives a concrete theoretical answer to a question that had previously been handled mostly by intuition: whether Wasserstein distances can act as a substitute for divergence measures like KL-divergence when comparing densities that share the same support. It shows they can, and it quantifies how rate and support information combine in finite samples — a regime that matters in practice, because real datasets are never infinite. Because the one-dimensional Wasserstein distance is the building block of the sliced Wasserstein distance, the results extend to the multidimensional settings where sliced Wasserstein distances are increasingly used.
Real-world applications suggested by the paper:
- Neuroscience: Decoding stimulus type from retinal ganglion cell spike trains and embedding human hippocampal spike trains for memory-task analysis.
- Brain-computer interfaces and neural signal interpretation: Distinguishing rapid rate changes from gradual temporal drift in recordings across electrodes.
- Molecular biology and protein structure: Building amino acid embeddings from contact frequency distributions that reflect hydrophobicity and long-range contact propensities.
- Generative modeling and domain adaptation: Where Wasserstein losses are already used as alternatives to KL-based objectives because they are less sensitive to support mismatch.
Industry relevance: Any pipeline that compares distributions over time or space — neural recordings, event logs, sensor streams, or molecular structure data — can use the transport-cost features described here. The paper provides open code at https://github.com/cheongjae/one-dim-wasserstein.
Future Directions
- Sliced and multidimensional Wasserstein distances. The paper sketches an illustrative example integrating sliced Wasserstein distances (Appendix C) but leaves a full treatment to future work.
- General distributions beyond the Poisson setting. The authors note that the derivation extends to any case where the double Laplace transform of |m⁻¹(u) − n⁻¹(v)| u^(k−1) v^(l−1) is analytically tractable, which is a restricted class.
- Unequal sample sizes and general pairing indices. The main derivations assume identical sample sizes N; the text states that the derivations can be extended to different sample sizes and time-varying rates, but the analysis in the paper focuses on the matched-N case.
- Embedding-based conclusions in neuroscience. The authors explicitly state that the Isomap embeddings "do not, by themselves, offer definitive conclusions," leaving room for further validation of the observed structure in neural activity.
Target Audience
Readers best served by this paper are machine learning researchers working on optimal transport and distribution comparison, theoretical statisticians interested in finite-sample properties of Wasserstein distances, and computational neuroscientists or bioinformaticians who apply transport-based distances to spike trains or molecular structure data. Practitioners using sliced Wasserstein distances in generative models or representation learning will also find the one-dimensional analysis directly relevant. A working knowledge of probability distributions and basic optimal transport is needed to follow the derivations.
Authors’ abstract
Leveraging the Wasserstein distance -- a summation of sample-wise transport distances in data space -- is advantageous in many applications for measuring support differences between two underlying density functions. However, when supports significantly overlap while densities exhibit substantial pointwise differences, it remains unclear whether and how this transport information can accurately identify these differences, particularly their analytic characterization in finite-sample settings. We address this issue by conducting an analysis of the information processing capabilities of the one-dimensional Wasserstein distance with finite samples. By utilizing the Poisson process and isolating the rate factor, we demonstrate the capability of capturing the pointwise density difference with Wasserstein distances and how this information harmonizes with support differences. The analyzed properties are confirmed using neural spike train decoding and amino acid contact frequency data. The results reveal that the one-dimensional Wasserstein distance highlights meaningful density differences related to both rate and support.