Skip to content
AI.info

Research

TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation

Overview Research area: Unsupervised domain adaptation (UDA) for medical image segmentation, specifically cross-modality adaptation (CT to MRI and MRI to CT), using vision-language (text-prompt) repre

TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation
arXiv
2511.05782
Published
2025-11-08
Authors
Lalit Maurya, Honghai Liu, Reyer Zwiggelaar

AI summary

Overview

Research area: Unsupervised domain adaptation (UDA) for medical image segmentation, specifically cross-modality adaptation (CT to MRI and MRI to CT), using vision-language (text-prompt) representation learning.

Technical level: Advanced. The paper builds on DeepLabV2/ResNet segmentation backbones, adversarial domain adaptation, Transformer-based vision-language fusion, dynamic convolution, prototype alignment, and pretrained language encoders (CLIP or BioBERT).

Scope: The paper proposes TCSA-UDA, a framework that uses modality-aware textual prompts to guide domain-invariant visual feature learning, and evaluates it on cardiac (MMWHS), abdominal (CHAOS/CT), and brain tumor (BraTS 2018) cross-modality segmentation benchmarks.

What This Paper Is About

Models trained to segment anatomical structures on one imaging modality (for example CT) often fail when applied to another modality (for example MRI), because scanners, protocols, and institutions introduce domain shift. Since annotating every new domain is impractical, the authors want a model that adapts from a labeled source modality to an unlabeled target modality without target annotations. Their key idea is that anatomical class names ("left atrium," "tumor") are inherently domain-invariant, so language can act as a semantic anchor while the prompt still carries modality context such as CT or MRI.

Key Contributions

  1. TCSA-UDA framework: A text-driven cross-semantic adaptation framework for unsupervised cross-modality medical image segmentation, which uses modality-aware textual prompting to guide domain-invariant visual representation learning.
  2. VLCoL (vision-language covariance cosine loss): A loss that aligns inter-class visual feature relationships (a covariance matrix over class prototypes) with text-derived anatomical semantic relationships (a covariance matrix over class text embeddings), using cosine distance between the two matrices.
  3. Modality-aware semantics injection plus class-level adaptation: Integration of modality-aware textual prompting, Transformer-based visual-language fusion, and dynamic convolution, combined with a prototype alignment module that reduces residual source-target class-level discrepancies.
  4. Validation across three clinical tasks: Experiments on cardiac, abdominal organ, and brain tumor cross-modality segmentation, with ablation studies to analyze the contribution of the key components (the ablation results are not included in the provided content).

Main Findings

  • MMWHS cardiac, MRI→CT (Table 1): Ours reaches an average Dice of 82.4 with average ASD 6.9, which is the highest average Dice among the listed domain adaptation methods. For comparison, MA-UDA reaches 81.1, SECASA 80.9, ADR 80.2, DAFormer 77.1, CRST 75.5, AdvEnt 75.0, EBM 74.9, SIFA V2 74.1, AdaOutput 69.6, PnP-AdaNet 63.9, and CycleGAN 57.6.
  • MMWHS cardiac, CT→MRI (Table 1): Ours reaches an average Dice of 71.6 with average ASD 5.1, again the highest average Dice listed. SECASA reaches 69.9, IPLC+ 68.9, MA-UDA 68.7, IPLC 68.1, SFDA-CCRC 66.9, ADR 66.3, DAFormer 65.9, CRST 64.8, EBM 64.0, SIFA V2 63.4, AdvEnt 63.2, AdaOutput 63.2, PnP-AdaNet 54.3, and CycleGAN 50.7.
  • Surface distance is not uniformly best: On MMWHS MRI→CT, the proposed method's average ASD of 6.9 is higher (worse) than SECASA (5.4), ADR (5.1), MA-UDA (5.6), SFDA-CCRC (5.8), IPLC+ (5.9), and CRST (6.8). On CT→MRI, its average ASD of 5.1 is higher (worse) than SFDA-CCRC (4.2), SECASA (4.3), and IPLC+ (4.6).
  • Supervised upper bound and no-adaptation floor on MMWHS: Fully supervised training reaches average Dice 90.4 and ASD 2.5 for MRI→CT, and 85.1 and 2.2 for CT→MRI. Without adaptation, performance collapses to average Dice 23.3 and ASD 22.6 (MRI→CT) and average Dice 20.4 and ASD 17.9 (CT→MRI).
  • Abdominal MRI→CT (Table 2): Ours reports average Dice 83.41 ± 0.30 and average ASD 0.71 ± 0.07, versus SIFA 81.52 ± 0.35, IPLC+ 80.77 ± 0.17, IPLC 80.29 ± 0.24, SECASA 78.62 ± 0.27, AdaOutput 78.57 ± 0.54, and AdvEnt 78.50 ± 0.61. Supervised training gives 88.64 ± 1.13 Dice and 0.54 ± 0.04 ASD.
  • Abdominal CT→MRI (Table 3): Ours reports average Dice 83.11 ± 0.29 and average ASD 1.06 ± 0.06, versus SIFA 81.37 ± 0.25, IPLC+ 80.90 ± 0.14, IPLC 80.73 ± 0.15, SECASA 78.84 ± 0.64, AdvEnt 74.54 ± 0.37, and AdaOutput 73.94 ± 0.61. Supervised training gives 87.92 ± 0.65 Dice and 0.92 ± 0.07 ASD.
  • Per-organ behavior on abdominal CT→MRI: The largest gain cited in the table is the Spleen, where Ours reports 77.30 ± 0.49 Dice versus AdaOutput 58.25 ± 1.06 and AdvEnt 57.62 ± 1.12.
  • Training objective: The total loss combines source segmentation loss (cross-entropy plus Dice), adversarial entropy alignment, the VLCoL covariance loss, and prototype alignment, weighted by λ1 = 0.003, λ2 = 1.0, and λ3 = 0.1.
  • Brain tumor and ablation results: BraTS 2018 experiments are described in the dataset section (bidirectional FLAIR↔T2 adaptation on the 75-subject LGG cohort, plus a 105-of-210 subject HGG subset), and ablation studies are listed as a contribution, but the corresponding result tables and numbers are not included in the provided content.

Methodology in Plain English

The framework adapts a segmentation network from a labeled source modality to an unlabeled target modality using four cooperating pieces.

  1. Modality-aware prompts. For every anatomical class, the authors build a sentence such as "A {Dataset} {CT/MRI} imaging of a [CLS]", where the dataset type and imaging modality provide context and the placeholder is filled with the class name (for example, myocardium of the left ventricle). These prompts are encoded by a shared pretrained language model — CLIP or BioBERT — so that all classes land in one common semantic space. Because the same encoder and the same class words are used for both domains, the resulting semantics are domain-invariant, while the modality token keeps modality context.

  2. Fusing text into vision. A DeepLabV2 encoder with a ResNet backbone produces high-level features. These are pooled into a single global visual vector, repeated once per class, and concatenated with each class's text embedding; a linear layer plus ReLU produces class-conditioned query vectors. A Transformer multi-head attention block lets these queries attend to the global visual feature (used as both key and value), and an MLP refines the result. The fused representation is then averaged over classes and passed through an MLP controller that generates per-sample dynamic convolution weights and biases, which modulate the segmentation features before the segmentation head. In effect, the text decides how the visual features should be reshaped for this particular modality and anatomy.

  3. VLCoL: matching relationships, not just features. For each class present, the encoder's pixel features are averaged under the (downsampled) ground-truth mask, and a memory bank with exponential decay smooths these class features across batches. The authors then compute a class-by-class covariance matrix of the visual class features and a matching covariance matrix of the text embeddings, and minimize the cosine distance between the two matrices. This forces the pattern of "which classes look similar to which" in the image encoder to mirror the pattern encoded in language — a semantic prior that does not depend on matching source and target distributions directly.

  4. Adversarial alignment and prototype alignment. Two GAN-style discriminators (a main one and an auxiliary one) receive entropy maps derived from the segmentation probabilities of source and target images; the discriminator is trained to tell them apart, and the segmentation network is trained to fool it, which pushes toward structurally consistent predictions across domains. Separately, class prototypes are computed by averaging penultimate-layer features per class — using ground-truth labels on the source and argmax pseudo-labels on the target — maintained as exponential moving averages, and pulled together with a squared L2 distance.

Training used a batch size of 4, SGD for the segmentation model (learning rate 2.5×10⁻⁴, momentum 0.9, weight decay 5×10⁻⁴), Adam for the discriminator and text fusion module (learning rate 1×10⁻⁴), images resized to 256×256, random scaling, rotation, and intensity variation as augmentation, a prototype momentum of 0.01, and a single NVIDIA A100 GPU with 40 GB. Performance was measured with Dice (higher is better) and Average Symmetric Surface Distance (lower is better), averaged over multiple model initializations, with the MMWHS partition following SIFA V2 and the abdominal and BraTS experiments using four random subject-level 80/20 splits with fixed seeds.

Why This Matters

Impact on research. Most existing UDA methods for medical segmentation align visual distributions through adversarial training, distance metrics such as MMD, or self-training with pseudo-labels, and the paper argues these pipelines lack explicit semantic consistency across heterogeneous modalities. This work shows that language-derived inter-class structure can serve as a training-time semantic prior rather than only as an inference-time prompt, which is a different use of vision-language models than prior work on open-vocabulary segmentation, prompt adaptation, or supervised single-domain medical segmentation.

Real-world applications (potential, as motivated by the paper's clinical framing of diagnosis and surgical planning):

  • Deploying a cardiac segmentation model trained on CT volumes at a site that acquires MRI, or vice versa, without re-annotating data at the new site.
  • Abdominal organ segmentation (spleen, right kidney, left kidney, liver) transferred between abdominal CT and T2-SPIR MRI acquisitions.
  • Brain tumor segmentation transferred between MRI contrasts, specifically FLAIR and T2.
  • Reducing annotation burden when onboarding new scanners, protocols, or institutions, by reusing labeled data from an existing modality.

Industry relevance. The setting mirrors commercial and clinical realities — heterogeneous scanners, protocols, and institutions — so a method that adapts without target annotations is directly relevant to medical imaging software vendors and clinical imaging workflows. Note that no deployment, regulatory, or cost figures are reported in this paper; relevance comes from the cross-modality adaptation setting.

Future Directions

  • Report and stress-test the brain tumor results and ablations. The paper describes BraTS 2018 FLAIR↔T2 experiments for LGG and an HGG subset and claims ablation studies, but these results are not present in the available content; their publication would clarify how much each component (VLCoL, prototype alignment, adversarial loss, dynamic convolution) contributes.
  • Resolve the Dice-versus-surface-distance trade-off. The method leads on average Dice but trails several baselines on average ASD in the MMWHS cardiac experiments, which raises the question of whether boundary accuracy needs a dedicated constraint.
  • Reduce the gap to supervised training. On MMWHS, supervised average Dice is 90.4 (MRI→CT) and 85.1 (CT→MRI) versus 82.4 and 71.6 for TCSA-UDA; on the abdominal tasks the gap is smaller (for example 88.64 versus 83.41 for MRI→CT), so the cross-modality cardiac setting remains far from the supervised ceiling.
  • Reduce dependence on the text encoder choice. Prompts are encoded with CLIP or BioBERT, and modality-aware prompt templates must be hand-written per dataset and modality; automating prompt construction and testing sensitivity to the language model choice are open questions the paper does not resolve.

Target Audience

Researchers and graduate students working on unsupervised domain adaptation, cross-modality medical image segmentation, and vision-language models in medical imaging. It is also relevant to applied practitioners in medical imaging who need segmentation models to transfer between CT and MRI acquisitions across sites, though the density of loss formulations and architecture details makes it best suited to readers with background in deep segmentation and adversarial adaptation.

Authors’ abstract

Unsupervised domain adaptation (UDA) for medical image segmentation remains challenging due to substantial domain shifts across imaging modalities, such as CT and MRI. Although recent vision-language representation learning methods have shown promise in medical image analysis, their role in cross-modality UDA segmentation remains underexplored. To address this problem, we propose TCSA-UDA, a Text-driven Cross-Semantic Alignment framework that uses modality-aware textual prompting to guide domain-invariant visual representation learning. Specifically, we introduce a vision-language covariance cosine loss (VLCoL) that aligns inter-class visual feature relationships with text-derived semantic relationships, encouraging the image encoder to learn semantically structured and modality-robust representations. In addition, we incorporate a prototype alignment module to reduce residual class-level discrepancies between source and target domains by aligning high-level class prototypes. Extensive experiments on cross-modality cardiac, abdominal, and brain tumor segmentation benchmarks demonstrate that TCSA-UDA consistently improves adaptation performance and outperforms state-of-the-art UDA methods. These results highlight the potential of language-driven semantic guidance for domain-adaptive medical image segmentation. The code is available at https://github.com/lalitmaurya47/TCSA_UDA

Read the original paper