Research
Leveraging Persistence Image to Enhance Robustness and Performance in Curvilinear Structure Segmentation
Overview Research area: Medical image analysis / computer vision — specifically curvilinear (vessel-like) structure segmentation using topological data analysis. Technical level: Advanced. The paper a

- arXiv
- 2601.18045
- Published
- 2026-01-25
- Authors
- Zhuangzhi Gao, Feixiang Zhou, He Zhao, Xiuju Chen, Xiaoxin Li, Qinkai Yu, Yitian Zhao, Alena Shantsila, Gregory Y. H. Lip, Eduard Shantsila, Yalin Zheng
AI summary
Overview
Research area: Medical image analysis / computer vision — specifically curvilinear (vessel-like) structure segmentation using topological data analysis.
Technical level: Advanced. The paper assumes familiarity with persistent homology, persistence diagrams, persistence images, Betti numbers, and standard segmentation backbones.
Scope: The paper proposes a framework (PIs-Regressor + Topology SegNet) that learns persistence images directly from raw images and fuses them into a segmentation network, evaluated on three curvilinear benchmarks plus a perturbed robustness benchmark.
What This Paper Is About
Segmenting curvilinear structures such as retinal vessels requires both pixel-level accuracy and topological correctness, because a segmentation can score well on Dice while still containing broken vessels or missing branches. Topology-aware methods usually try to fix this with handcrafted loss functions, but persistence diagrams themselves are discrete, non-differentiable, and expensive to compute, so those losses generalize poorly. This paper instead learns a differentiable, fixed-size topological representation — the persistence image — from data and inserts it directly into the segmentation architecture rather than into the loss.
Key Contributions
- PIs-Regressor: a compact network that directly learns an approximation of the persistence image (PI) from raw input images, avoiding the non-differentiability and computational cost of computing persistence diagrams on the fly.
- Topology SegNet: a segmentation network that fuses topological features with image features at both the downsampling and upsampling stages — the authors state this is the first work to incorporate topological features into classical segmentation networks to enhance performance.
- Architecture-level topology instead of loss-level topology: rather than relying on handcrafted topological losses, the framework embeds topology into the network structure, and the design can be combined with existing topology-based losses (PH Loss, clDice, cbDice) for further gains.
- Demonstrated robustness and state-of-the-art results on three curvilinear benchmarks (DRIVE, ER, and a private OPTOS ultra-widefield retinal dataset) in both pixel-level accuracy and topological fidelity, plus a perturbation study on DRIVE.
Main Findings
- DRIVE pixel and topological gains: On DRIVE, the model with the traditional CE loss improves the Dice coefficient by 1.18 and reduces topological error (measured by β₀) by 90.4 compared to a standard U-Net with CE loss. Concrete values: Ours + CE Loss reaches Dice 81.91, clDice 82.26, MIoU 82.12, β₀ 126.8, β₁ 22.65, versus CE Loss at 80.73 / 81.00 / 81.13 / 217.2 / 23.2.
- Best DRIVE configuration: Ours+CE+Dice Loss gives Dice 82.12, clDice 82.56, MIoU 82.30, β₀ 101.25, β₁ 24.95.
- ER dataset — the largest connectivity gains: The best-performing baseline is U-Net with clDice loss. Ours + clDice achieves a 2.88-point improvement in clDice along with reductions of 53.5 in β₀ and 46.02 in β₁. Absolute values: Ours + clDice reaches Dice 84.52, clDice 93.48, MIoU 79.80, β₀ 30.70, β₁ 31.83, compared with clDice baseline at 84.23 / 90.60 / 79.62 / 184.2 / 77.85.
- Private OPTOS dataset: Among baselines, U-Net with cbDice loss performs best. Using the authors' model with CE and Dice loss improves Dice by 1.84, clDice by 1.69, and mIoU by 0.92, while reducing β₀ by 4.23 and β₁ by 5.55 relative to that best baseline. Absolute values: Ours+CE+Dice Loss at Dice 76.99, clDice 82.77, MIoU 87.72, β₀ 4.278, β₁ 0.464; cbDice baseline at 75.15 / 81.08 / 86.80 / 8.503 / 6.010.
- Robustness under overexposure (DRIVE): Traditional segmentation (U-Net + CE loss) shows large degradation; β₀ and β₁ increased by 1039.9 and 43.6 respectively, indicating more fragmentation and noise. The proposed method saw a much smaller increase in β₀ (76.9) and a decrease in β₁ (-0.8). It outperformed U-Net with cbDice on Dice, clDice, and MIoU by 0.72, 0.06, and 0.68 under overexposure.
- Robustness under blur, low contrast, underexposure: Ours+CE+Dice Loss on Blur DRIVE reaches Dice 66.48, clDice 65.83, MIoU 71.30, β₀ 50.00, β₁ 46.45 (CE Loss: 64.82 / 56.52 / 70.37 / 95.00 / 54.25). On Contrast DRIVE: 82.16 / 82.28 / 82.38 / 96.5 / 25.85. On Underexposed DRIVE: 81.03 / 80.71 / 81.47 / 94.5 / 28.25.
- Perturbation protocol: The robustness comparison uses models pretrained on the unperturbed DRIVE dataset and tested on four perturbations: blur, low contrast, underexposure, and overexposure.
- Combination capability: The PI-based design is reported as flexible and can be combined with PH Loss, clDice, and cbDice, generally improving over the corresponding standalone baselines across datasets.
- Notable exception: On the private dataset, cbDice alone reports β₁ of 6.010 while Ours + cbDice reports β₁ of 7.289, and on DRIVE, clDice's β₁ (26.2) is lower than Ours + clDice (23.7 is lower — actually Ours is lower here). The paper does not claim uniform improvement on every single metric/dataset pairing.
Methodology in Plain English
The pipeline starts from a segmentation map treated as a sparse point cloud. As a scale parameter grows, nearby points merge, forming connected components and loops. Persistent homology tracks when these features are born and when they die, producing a persistence diagram — but a diagram is a discrete multiset and cannot be fed to a neural network or differentiated through.
The authors' workaround is the persistence image: diagram points are transformed from (birth, death) to (birth, death − birth) coordinates, weighted by a function and smoothed with Gaussians to form a continuous persistence surface, which is then discretized into a grid of cells by integrating over each cell. That gives a fixed-size image-like representation of topology.
Rather than computing this image from the ground truth and back-propagating through a non-differentiable pipeline, the PIs-Regressor learns to predict it directly from the raw input image. It uses a pre-trained ResNet-50 encoder, a DeConv Block decoder, a 1×1 convolution and Sigmoid activation, producing an output of shape b×1×h×w, trained with Mean Squared Error loss. The paper notes that because the ground-truth PI is Gaussian-smoothed, it emphasizes global topological structure (vessel connectivity and circularity) rather than local detail, which gives inherent tolerance to local prediction error.
The predicted persistence image is then repeated and upsampled to h×w×3 and concatenated with the transformed input image, producing a tensor of shape b×h×w×(c+3). This fused tensor enters Topology SegNet. Inside, at every downsampling step the max-pooled features are fused with topological features resized to match, and the result passes to the next downsampling module. The same fusion is applied during upsampling, but with features obtained after deconvolution-based upsampling.
Training uses two subnetworks with different optimizers: PIs-Regressor uses Adam at 10⁻³, while Topology SegNet uses SGD at 10⁻² with a scheduler and 0.9 decay. Training runs 500 epochs with batch size 4, implemented in PyTorch, with Betti numbers computed using GUDHI. The segmentation loss is α·Dice + (1−α)·CE with α set to 0.5 for simplicity. Evaluation metrics are IoU (reported as MIoU), clDice, Dice coefficient, and Betti Error using β₀ and β₁.
Why This Matters
Impact on research: The paper argues for a shift from topology-as-a-loss-term to topology-as-an-architecture-component. Because handcrafted topological losses are task-specific and can introduce artifacts (the authors point out that skeletonization in clDice can introduce errors that compromise segmentation quality), learning a differentiable topological representation from data offers a more general and transferable route. It also reframes topology as more than a connectivity constraint — the authors stress that the multifaceted properties of topological features extend beyond merely capturing connectivity.
Real-world applications:
- Retinal vessel analysis for early diagnosis and treatment planning in retinal and cardiovascular disease, where broken vessels or missing small branches can cause clinical misjudgment.
- Clinical deployment under imperfect imaging, e.g., overexposed, blurred, or low-contrast fundus photographs, where the paper shows large robustness advantages over standard training.
- Ultra-widefield retinal imaging (the private OPTOS dataset), where vessel morphology is complex and structural continuity is easily lost.
- General curvilinear structure segmentation beyond vessels, since the approach is not tied to a specific anatomy or loss function.
Industry relevance: Medical imaging software and screening pipelines depend on measurements derived from segmented vessels; topological errors propagate into downstream morphological measurements. A module that can be bolted onto classical segmentation backbones and combined with existing topological losses is attractive for regulated clinical products where robustness to acquisition variability matters.
Future Directions
- The paper releases code at https://github.com/NatsuGao7/TopoUnet.git, but does not lay out an explicit future-work agenda; the following follow logically from the reported results.
- Resolving metric inconsistencies: in several dataset/method pairings the proposed model does not beat every baseline on every metric (for example, β₁ on the private dataset with cbDice, and β₁ on DRIVE with clDice). Understanding when fusing PI helps versus hurts is an open question the paper does not resolve.
- Extending beyond H₁: the analysis focuses on 1-dimensional homology (loops) and notes that H₀ and H₁ are typically used in 2D images. Whether higher-dimensional or alternative homology features add value is not explored.
- Generalization scope: evaluation covers DRIVE, ER, and one private retinal dataset, all vascular. Performance on other curvilinear structures, imaging modalities, or 3D data is not reported.
- Coupling with more topology-based losses: the paper states the design can be seamlessly combined with other topology-based methods to further enhance performance, but only demonstrates this with PH Loss, clDice, and cbDice.
Target Audience
Researchers and graduate students working on medical image segmentation, topological data analysis in deep learning, and retinal vessel analysis. It is also relevant to computer vision practitioners interested in differentiable topological representations, and to clinical/industrial teams building robust vessel-segmentation pipelines who already understand standard U-Net-style architectures and are willing to engage with persistent homology notation. Beginners would struggle with the preliminaries section, which assumes prior exposure to homology groups, persistence barcodes, and persistence diagrams.
Authors’ abstract
Segmenting curvilinear structures in medical images is essential for analyzing morphological patterns in clinical applications. Integrating topological properties, such as connectivity, improves segmentation accuracy and consistency. However, extracting and embedding such properties - especially from Persistence Diagrams (PD) - is challenging due to their non-differentiability and computational cost. Existing approaches mostly encode topology through handcrafted loss functions, which generalize poorly across tasks. In this paper, we propose PIs-Regressor, a simple yet effective module that learns persistence image (PI) - finite, differentiable representations of topological features - directly from data. Together with Topology SegNet, which fuses these features in both downsampling and upsampling stages, our framework integrates topology into the network architecture itself rather than auxiliary losses. Unlike existing methods that depend heavily on handcrafted loss functions, our approach directly incorporates topological information into the network structure, leading to more robust segmentation. Our design is flexible and can be seamlessly combined with other topology-based methods to further enhance segmentation performance. Experimental results show that integrating topological features enhances model robustness, effectively handling challenges like overexposure and blurring in medical imaging. Our approach on three curvilinear benchmarks demonstrate state-of-the-art performance in both pixel-level accuracy and topological fidelity.