Skip to content
AI.info

Research

GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs

GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs Overview Research area: Graph machine learning — specifically text-attributed graphs (TAGs), graph con

GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs
arXiv
2511.16778
Published
2025-11-20
Authors
Yating Ren, Yikun Ban, Huobin Tan

AI summary

GCL-OT: Graph Contrastive Learning with Optimal Transport for Heterophilic Text-Attributed Graphs

Overview

Research area: Graph machine learning — specifically text-attributed graphs (TAGs), graph contrastive learning (GCL), and optimal transport (OT) applied to heterophilic graphs.

Technical level: Advanced. The paper assumes familiarity with graph neural networks, contrastive learning objectives (InfoNCE), optimal transport / Sinkhorn algorithms, and mutual information bounds.

Scope: The paper introduces GCL-OT, a contrastive framework that uses optimal transport to softly and bidirectionally align structural and textual representations in text-attributed graphs, with tailored mechanisms for three identified forms of heterophily, evaluated on nine benchmark datasets.

What This Paper Is About

Text-attributed graphs connect text-bearing nodes (papers, products, web pages) via edges, but many real graphs are heterophilic — connected nodes often belong to different classes, so their texts and structures disagree. Existing structure–text contrastive methods assume homophily and rely on hard, one-to-one pairing objectives, which misalign representations when semantic correlations are mixed, noisy, or missing. GCL-OT's goal is to replace hard matching with flexible many-to-many alignment driven by optimal transport, so that structural and textual views can be reconciled even when a node's text aligns with only some neighbors, no neighbors, or with unconnected but semantically similar nodes.

Key Contributions

  1. First integration of optimal transport into graph contrastive learning for heterophilic text-attributed graphs, enabling flexible, bidirectional alignment between structural and textual views rather than one-sided greedy matching.
  2. Three mechanisms matched to three granularities of heterophily: a RealSoftMax-based similarity estimator for partial heterophily, a prompt-based filter for complete heterophily, and OT-guided soft supervision for latent homophily.
  3. Theoretical analysis showing GCL-OT yields tighter mutual information lower bounds than InfoNCE and a tighter upper bound on downstream Bayes error.
  4. Extensive experiments on nine benchmarks spanning homophilic and heterophilic regimes, plus robustness, ablation, efficiency, PLM, and hyperparameter studies.

Main Findings

  • Multi-granular heterophily is identified and characterized: complete heterophily, partial heterophily, and latent homophily. The paper argues these induce many-to-many (N:N) alignment between structural and textual neighborhoods, which DGI's near-N:1 pairing and InfoNCE's approximately 1:1 correspondence cannot capture.
  • Supervised node classification: GCL-OT variants lead on all nine datasets. Ours-SAGE reaches 93.73±1.9 (Cora), 96.62±0.6 (PubMed), 81.73±0.5 (Products), 78.13±0.2 (ArXiv), 66.28±0.5 (Amazon), 82.72±2.8 (ArXiv23), 89.26±4.9 (Wisconsin), 88.21±6.0 (Cornell), and 90.01±4.4 (Texas). Comparison points include TAPE-RevGAT at 47.22±1.0 on Amazon and 92.80±2.8 on Cora, and ENGINE-LLAMA at 54.60±0.9 on Amazon.
  • Largest gains appear on heterophilic data: on Amazon, where node homophily is 37.57, Ours-GCN scores 66.40±0.4 versus 46.39±2.4 for TAPE-SAGE and 47.22±1.0 for TAPE-RevGAT. Texas (homophily 10.68) reaches 89.47±3.4 with Ours-GCN and 90.01±4.4 with Ours-SAGE.
  • Unsupervised setting: with a frozen model and a one-layer linear classifier on 60/20/20 splits, Ours-GCN scores 72.83±7.1 (Wisconsin), 64.10±4.4 (Cornell), and 73.68±3.4 (Texas), exceeding HeterGCL (71.93±4.2, 64.02±4.8, 72.79±2.5), PolyGCL (70.24±4.3, 62.45±5.5, 66.77±5.5), Congrat, GRACE, and DGI.
  • Robustness to structural perturbation: under edge perturbation from -500 (removal) to +500 (addition) in steps of 50, vanilla GCN and GAT drop roughly 50% and 75% at the -500 worst condition, while GCL-OT with GCN and GCL-OT with GAT achieve relative gains of 24.74% and 79.44% respectively.
  • Robustness to text perturbation: across 0% to 100% of nodes with word perturbation from -100% to +100% in steps of 10%, DistilBERT degrades significantly whereas GCL-OT retains nearly 85%, showing a smoother performance curve.
  • Case study vs. InfoNCE: replacing the GCL-OT loss with InfoNCE yields substantial drops for GCL-OT, with the largest improvements under mixed-label neighborhoods and strong semantic heterophily; the paper notes benefits may be limited where purely heterophilic neighborhoods and unlinked similar pairs are abundant.
  • Ablation: removing the contrastive loss degrades performance significantly; removing the alignment loss causes large drops on heterophilic datasets such as Texas and Cornell; removing the latent homophily mining loss hurts datasets with weak structural signals such as Amazon.
  • Parameter sensitivity: λ and β were varied over {0.0, 0.1, ..., 1.0} and LRSinkhorn iterations over {10, 20, 30, 40} on Cora with GCN. Suitable values help; λ influences stability as reflected in standard deviation, with only mild fluctuations beyond these values.
  • PLM substitution: swapping DistilBERT for DeBERTa or RoBERTa yields comparable or slightly better accuracy on most datasets — DeBERTa at 89.16±4.7 (Wisconsin), 87.35±6.2 (Cornell), 87.30±11.8 (Texas); RoBERTa at 86.42±6.6, 88.21±3.9, and 86.82±14.6.
  • Efficiency: reported per-epoch complexity is O(|E|D + NW²D + NrD + Nr). The text-encoder term O(NW²D) dominates in typical benchmarks, so cost scales nearly linearly with node count and feature dimension and quadratically with text length. GCL-OT with GCN reportedly achieves the highest accuracy at a relatively short training time.
  • Theoretical results: Propositions 1 and 2 state MI(H^ζ, H^t) ≥ log N − L_MHA ≥ log N − L_InfoNCE and MI(H^ζ, H^t) ≥ log N − L_LHM ≥ log N − L_InfoNCE. A remark gives L_GCL-OT ≥ 4 log r − 2(JS^ζ + JS^t). Theorem 1 covers the fused embedding under a full-column-rank linear fusion assumption.
  • Visualization: t-SNE plots on Cora and Texas show GCL-OT producing more semantically coherent and better-separated clusters than baselines.

Methodology in Plain English

The framework has four moving parts.

  1. Feature encoding. Each node's text is enriched by prompting a frozen LLM (GPT-3.5) and concatenating the answer with the original text. A partially frozen DistilBERT (six layers) then produces both token-level embeddings and a sentence-level embedding. In parallel, a GNN (GCN, GAT, or GraphSAGE, each with two layers) produces structural embeddings, from which a neighborhood-level matrix is formed for each node.
  2. Soft similarity instead of averaging. Comparing a node's neighbors to another node's words is done with RealSoftMax, a smooth operator that interpolates between mean (β → ∞) and maximum (β → 0). Two symmetric terms are averaged — one emphasizing the most relevant words per neighbor and the other the reverse — so key cross-view interactions are amplified while background noise is downweighted.
  3. Filtering unalignable content. The similarity matrix is augmented with a learnable prompt row and column, giving an (N+1) × (N+1) matrix. Embeddings whose best similarity falls below the prompt entry get routed to the prompt instead of being forced into a bad match. The alignment itself is a Sinkhorn optimal transport problem with an entropy regularizer, solved via a non-negative low-rank factorization, with uniform marginals on both sides. The resulting transport plan and similarity feed a bi-directional contrastive loss that keeps the top r scoring negatives per row.
  4. Discovering hidden neighbors. The OT assignment matrix is added to the identity matrix (self-positives) to form a soft target distribution, and a latent homophily mining loss trains the model to match that distribution — effectively pulling together semantically similar but unconnected nodes instead of treating them as negatives.

Training combines a cross-entropy node-classification loss with the contrastive term, weighted by λ. Experiments used one NVIDIA 4090 GPU and three NVIDIA 3080 GPUs; text and graph encoder dimensions were both 768, learning rates 0.01 and 2e-5, weight decay 0.0005, dropout from 0 to 0.8, batch size 512, and results are averages over 10 random seeds. ArXiv and Products used standard splits; the other datasets used random 60/20/20 splits.

Why This Matters

Research impact. The paper reframes structure–text alignment in heterophilic TAGs as a transport problem rather than a hard matching problem, and it introduces a vocabulary (partial, complete, latent heterophily) for describing why alignment fails. The mutual information and Bayes error analysis gives a theoretical handle for why soft, many-to-many objectives can beat InfoNCE-style pairing.

Real-world applications:

  • Academic citation networks and preprint archives, where linked papers are often on different topics.
  • E-commerce recommendation, where co-purchase edges can be arbitrary or accidental and product descriptions vary wildly in relevance.
  • Web hyperlink graphs, where pages link across disparate domains.
  • Social and dating networks, where the "opposites attract" principle produces edges between dissimilar profiles.

Industry relevance. The framework is architecture-agnostic — it plugs into GCN, GAT, or GraphSAGE and works with DistilBERT-level encoders, so practitioners do not need the largest available language models. It also degrades gracefully under edge and text noise, which matters in production graphs where links and descriptions are messy, and it supports an unsupervised regime useful when labels are scarce.

Future Directions

  • Scaling the transport solver. The reported per-epoch cost is dominated by O(NW²D), so extending to very large graphs and longer texts remains an open engineering question.
  • Adaptive hyperparameter control. λ, β, and the number of LRSinkhorn iterations are currently selected through grid search over {0.0, 0.1, ..., 1.0} and {10, 20, 30, 40}; learning or scheduling them could remove a costly tuning step.
  • When soft alignment does not help. The case study notes that benefits may be limited when purely heterophilic neighborhoods and unlinked similar pairs dominate; characterizing that failure regime is a natural follow-up.
  • Beyond node classification. The paper evaluates node classification (supervised and unsupervised) and t-SNE visualization; extending the OT-guided supervision to link prediction, clustering, and anomaly detection is left open.

Target Audience

Researchers and graduate students working on graph representation learning, graph contrastive learning, or optimal transport applied to structured data, along with practitioners building text-and-graph systems for recommendation, search, or citation analysis who face heterophilic or noisy graphs. Readers should already be comfortable with GNN message passing, InfoNCE-style objectives, and the basics of Sinkhorn-based optimal transport.

Authors’ abstract

Recently, structure-text contrastive learning has shown promising performance on text-attributed graphs by leveraging the complementary strengths of graph neural networks and language models. However, existing methods typically rely on homophily assumptions in similarity estimation and hard optimization objectives, which limit their applicability to heterophilic graphs. Although existing methods can mitigate heterophily through structural adjustments or neighbor aggregation, they usually treat textual embeddings as static targets, leading to suboptimal alignment. In this work, we identify multi-granular heterophily in text-attributed graphs, including complete heterophily, partial heterophily, and latent homophily, which makes structure-text alignment particularly challenging due to mixed, noisy, and missing semantic correlations. To achieve flexible and bidirectional alignment, we propose GCL-OT, a novel graph contrastive learning framework with optimal transport, equipped with tailored mechanisms for each type of heterophily. Specifically, for partial heterophily, we design a RealSoftMax-based similarity estimator to emphasize key neighbor-word interactions while easing background noise. For complete heterophily, we introduce a prompt-based filter that adaptively excludes irrelevant noise during optimal transport alignment. Furthermore, we incorporate OT-guided soft supervision to uncover potential neighbors with similar semantics, enhancing the learning of latent homophily. Theoretical analysis shows that GCL-OT can improve the mutual information bound and Bayes error guarantees. Extensive experiments on nine benchmarks show that GCL-OT outperforms state-of-the-art methods, demonstrating its effectiveness and robustness.

Read the original paper