Research
Gromov-Wasserstein Distillation for Inductive Multi-View Embedding
Overview Research area: Machine learning — relational dimensionality reduction, optimal transport, and knowledge distillation. Specifically, Gromov–Wasserstein geometry applied to single-view and mult

- arXiv
- 2609.40047
- Published
- 2026-09-30
- Authors
- Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante
AI summary
Overview
Research area: Machine learning — relational dimensionality reduction, optimal transport, and knowledge distillation. Specifically, Gromov–Wasserstein geometry applied to single-view and multi-view embedding.
Technical level: Intermediate. The paper relies on optimal transport concepts (transport plans, barycentric projections, Gromov–Wasserstein discrepancy) that require some mathematical maturity, but its central idea is explained in accessible terms and the experimental protocol is standard.
Scope in one sentence: The paper turns a transductive Gromov–Wasserstein embedding method into an inductive neural mapping by distilling sample-aligned "barycentric" targets from a teacher into a student network.
What This Paper Is About
Gromov–Wasserstein multidimensional scaling (GW-MDS) produces low-dimensional coordinates that preserve the relational structure of data, but it can only produce coordinates for the samples it was trained on, and the correspondence between input samples and learned coordinates is mediated by an optimal transport plan rather than being fixed. The authors address this by using the teacher's transport plan to build sample-aligned target coordinates through a barycentric projection, then training a neural network to predict those targets from raw inputs, so unseen samples can be embedded with a single forward pass.
Key Contributions
-
A barycentric distillation mechanism that resolves the sample-correspondence ambiguity of GW embeddings, converting transductive representations into explicit inductive mappings. The paper proves (Proposition 3.1 / A.1) that the barycentric projection is the unique minimizer of a transport-weighted quadratic reconstruction cost, and that the cost for any candidate representation decomposes into the optimal cost plus a non-negative weighted squared deviation term.
-
Single-view and multi-view inductive formulations derived from three teacher types: GW-MDS for single-view data, Mean-GWMDS (which averages the relational matrices across views before computing the teacher) and Multi-GWMDS (which jointly minimizes view-dependent GW discrepancies), with a training-only projection-selection criterion for the Multi-GWMDS case.
-
A direct neural GW baseline — a network trained purely by minimizing a Gromov–Wasserstein objective, without barycentric supervision — used to isolate the contribution of the distillation step.
-
An evaluation on synthetic and real-world data across multiple relational geometries (Euclidean, geodesic, cosine), showing consistent improvements over direct neural GW training and competitive or superior performance relative to PCA-based inductive baselines.
Main Findings
-
Barycentric distillation beats direct neural GW on the synthetic S-curve. Under geodesic relations, Inductive GW-MDS reached Pearson 0.9980 and Spearman 0.9974, trustworthiness 0.9981, and stress 0.0337, compared with 0.7142, 0.7248, 0.9233, and 0.3905 for direct neural GW. Under cosine relations the gap was even wider: Pearson 0.9513 versus 0.1904, and stress 0.2644 versus 0.6648.
-
Under Euclidean relations on the S-curve, PCA is competitive. PCA achieved Pearson 0.8959, Spearman 0.9075, and stress 0.1971 versus 0.8869, 0.8911, and 0.2003 for Inductive GW-MDS, while Inductive GW-MDS obtained the highest trustworthiness (0.9300 versus PCA's 0.9149).
-
Inductive Mean-GWMDS achieves the best value on all four metrics under every relational geometry on ERA5. For geodesic relations it reached Pearson 0.8507, Spearman 0.8439, trustworthiness 0.9373, and stress 0.2603. Its advantage over concatenated PCA under Euclidean relations was modest — 0.0063 in Pearson correlation and 0.0118 in stress — and larger for cosine and geodesic dissimilarities.
-
On ERA5, view-wise performance is uneven. Under Euclidean geometry, Inductive Mean-GWMDS obtained Pearson correlations of 0.8995, 0.9232, and 0.8981 for temperature, dewpoint temperature, and surface pressure, but only 0.4203 for total precipitation; geodesic relations raised the precipitation correlation to 0.7021 while keeping the other three views above 0.87. Under cosine dissimilarity, surface pressure had Pearson 0.2667 but Spearman 0.6706.
-
On rMD17–Aspirin, Inductive Mean-GWMDS also leads. Cosine dissimilarity gave the highest global correlations, and the advantage over concatenated PCA was clearest there, with Pearson rising from 0.4222 to 0.5242; geodesic relations gave the highest trustworthiness and lowest stress, though the differences relative to Euclidean geometry were small.
-
Averaged multi-view results hide a strong asymmetry between molecular views. The selection criterion chose the interatomic-distance projection under all three geometries. Under cosine dissimilarity, Inductive Multi-GWMDS attained correlations of 0.8204 and 0.0669 with the interatomic-distance and force-derived structures, whereas Inductive Mean-GWMDS obtained 0.7740 and 0.2743.
-
Distillation preserves teacher geometry closely. On the training set, the difference in weighted Pearson correlation between each teacher and its student was at most 0.0197 for Mean-GWMDS and 0.0045 for the selected Multi-GWMDS projection.
-
A low GW loss does not imply good sample-indexed organization. Direct multi-view GW reached objective values nearly identical to the Multi-GWMDS teacher (0.00920 versus 0.00919 under cosine dissimilarity on rMD17–Aspirin), and attained a slightly lower geodesic objective on ERA5, yet produced only r = 0.3409 there, versus 0.7855 for Inductive Multi-GWMDS and 0.8507 for Inductive Mean-GWMDS.
-
Qualitative embeddings agree with the quantitative pattern. On ERA5, held-out locations were positioned coherently by latitude, with less mixing than in the direct multi-view GW representation; on rMD17–Aspirin, Inductive Mean-GWMDS preserved a branched consensus structure while the direct baseline was more diffuse.
Methodology in Plain English
The approach works in two stages. First, a transductive teacher — GW-MDS for single-view data, Mean-GWMDS or Multi-GWMDS for multi-view data — learns low-dimensional coordinates whose pairwise distances match the relational structure of the input. Because the correspondence between input samples and learned coordinates is described by an optimal transport plan rather than a direct index pairing, the raw teacher coordinates cannot be used as per-sample regression targets.
The fix is the barycentric projection: each sample's target is the transport-weighted average of the latent points associated with it, which produces a representation explicitly aligned with the input indices. A neural student then learns to map raw observations to these targets by minimizing a simple mean squared error over training samples. For multi-view data, Mean-GWMDS averages the views' relational matrices before computing the teacher, while Multi-GWMDS learns a shared support from all views jointly and generates one projection per view; a single projection is then chosen by comparing the Pearson correlation between each candidate projection's induced distance matrix and the training views' relational matrices, aggregating across views.
All relational matrices are normalized by their maximum entry, and the barycentric targets are standardized using training-set statistics only. The student is trained against the targets with a validation split and early stopping. At inference, unseen samples are embedded by a single forward pass — no new pairwise relational matrix is constructed and no GW problem is solved.
The empirical work covers a synthetic S-curve with 1,250 samples at noise level 0.05 (1,000 training, 250 test), plus two real datasets. ERA5 consists of hourly January 2024 observations over a 23 × 21 = 483 location grid with 744 time instants, giving four co-registered views built from 2-m temperature, 2-m dew-point temperature, surface pressure, and total precipitation, each a 483 × 744 matrix. rMD17–Aspirin uses 1,000 approximately equidistant conformations selected from a 100,000-conformation trajectory of a 21-atom molecule, with two rotation-invariant views: 210 pairwise interatomic distances and 231 upper-triangular entries of the force Gram matrix. Both real datasets were split 80% train / 20% test (386/97 ERA5 locations, 800/200 conformations), with 15% of each training set held out for validation. Evaluation used Pearson correlation, Spearman correlation, trustworthiness at k = 10, and scale-adjusted stress.
Why This Matters
Impact on research. The paper identifies a specific conceptual trap: minimizing a Gromov–Wasserstein objective can achieve structural agreement under an optimized coupling without recovering the sample-indexed correspondence that out-of-sample prediction requires. The barycentric projection is presented as a principled bridge between transductive GW embeddings and inductive neural mappings, and the decomposition result in Proposition 3.1 justifies it as the unique minimizer of a transport-weighted reconstruction cost. This gives practitioners a concrete reason to prefer distillation over direct GW-loss training when an inductive mapping is the goal.
Real-world applications (grounded in the data domains the paper evaluates):
- Climate and atmospheric science: embedding spatial locations from reanalysis fields such as temperature, dewpoint, surface pressure, and precipitation into a shared low-dimensional space that can be applied to new locations.
- Molecular simulation: representing conformations of small molecules using rotation-invariant descriptors such as interatomic distances and force Gram matrices, with the ability to embed newly generated conformations.
- Multi-sensor and multi-modal data fusion, where the same observations are described by heterogeneous measurements that cannot be compared coordinate-wise.
- Any setting requiring out-of-sample embedding from relational data where new pairwise matrices would be expensive to construct at inference time.
Industry relevance. The central practical claim is efficiency: once the student is trained, embedding a new sample costs one forward pass, with no relational-matrix construction and no GW optimization at inference. That matters for deployed systems where new data arrives continuously. The paper also states the method avoids the coordinate-comparability requirement of conventional multi-view approaches, which broadens the set of data sources that can be fused. The authors note that scalable teacher optimization and mini-batch variants remain open, so the current formulation is not yet positioned for very large-scale deployment.
Future Directions
- Scalable teacher optimization. The paper lists this explicitly as future work, since the transductive GW teacher currently solves a full-support transport problem over the training set.
- A clustering-oriented variant with mini-batch optimization. The authors propose adapting the framework toward clustering and equipping it with mini-batch training, which would relax the full-support uniform-weight setting used throughout.
- Broader evaluation. Larger datasets, repeated data splits, and additional multi-view settings are named as needed to test whether the observed advantages generalize beyond the three datasets studied.
- Handling partially incompatible views. Both real-data experiments reveal strong asymmetries between views — for example, the precipitation view on ERA5 and the force-derived view on rMD17–Aspirin — and the paper interprets this as evidence that some relational structures are only partially compatible within a shared low-dimensional representation. How best to balance such views remains an open question.
Target Audience
Researchers and graduate students working on optimal transport, Gromov–Wasserstein geometry, dimensionality reduction, and knowledge distillation will find the core contribution most directly useful, particularly the theoretical justification of barycentric targets. Practitioners in climate science and computational chemistry who need inductive embeddings of relational or multi-view data are the most likely applied audience, since the two real-world evaluations come from those domains. Readers without a background in optimal transport will need to work through the background section to follow the formulation, but the experimental comparison and the direct-versus-distilled framing are accessible without it.
Authors’ abstract
Gromov-Wasserstein multidimensional scaling (GW-MDS) learns low-dimensional representations from relational data but remains transductive, providing no explicit mapping for unseen samples. We introduce an inductive framework based on barycentric distillation. A GW-MDS teacher learns a latent support and an optimal transport plan from the training data, and barycentric projection converts the resulting coupling into sample-aligned targets. A neural student then learns an explicit out-of-sample mapping, avoiding additional relational-matrix construction and GW optimization at inference. We formulate the approach for single-view data and extend it to Mean-GWMDS and Multi-GWMDS teachers through consensus and selected-projection targets learned by a multi-view student with view-specific encoders. We also investigate a direct neural baseline trained solely with a GW objective. Experiments on synthetic and real-world data using Euclidean, geodesic, and cosine relations show that the distilled models preserve the teacher geometry on unseen samples and consistently outperform direct neural GW training in sample-indexed relational preservation. These results establish barycentric projection as an effective bridge between transductive GW embeddings and inductive neural mappings.