Research
MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal Prostate MRI Segmentation
Overview Research area: Medical computer vision — automated prostate MRI segmentation for longitudinal Active Surveillance of prostate cancer. Technical level: Advanced. The paper combines 3D medical

- arXiv
- 2510.17529
- Published
- 2025-10-20
- Authors
- Yovin Yahathugoda, Davide Prezzi, Patricia A. Gutierrez, Piyalitt Ittichaiwong, Vicky Goh, Sebastien Ourselin, Michela Antonelli
AI summary
Overview
Research area: Medical computer vision — automated prostate MRI segmentation for longitudinal Active Surveillance of prostate cancer.
Technical level: Advanced. The paper combines 3D medical image segmentation, cross-attention, state-space models (Mamba), and semi-supervised self-training.
Scope: The paper proposes MambaX-Net, a semi-supervised, dual-scan 3D segmentation architecture that segments the prostate at one MRI time point by using the MRI and segmentation mask from the previous time point, trained with pseudo-labels rather than expert annotations.
What This Paper Is About
Active Surveillance (AS) manages low- and intermediate-risk prostate cancer by monitoring disease over time with serial MRI instead of immediately treating it. Automating that monitoring requires accurate prostate segmentation, but most deep-learning segmentation models are trained on single-time-point data with expert-drawn labels, which does not match the longitudinal, label-scarce reality of AS. The paper's goal is a segmentation model built specifically for the longitudinal setting: one that can exploit a patient's earlier scan and its mask, and that can learn without a large supply of expert annotations.
Key Contributions
- MambaX-Net, a semi-supervised, dual-scan 3D segmentation architecture that produces the segmentation for time point t by using the MRI from time point t together with the MRI and corresponding segmentation mask from the previous time point.
- A Mamba-enhanced Cross-Attention Module, which integrates a Mamba block into cross-attention so the model can capture temporal evolution across scans and long-range spatial dependencies efficiently.
- A Shape Extractor Module, which encodes the previous time point's segmentation mask into a latent anatomical representation used for more refined prostate zone delineation.
- A semi-supervised self-training strategy that uses pseudo-labels generated by a pre-trained nnU-Net, allowing the model to learn effectively without expert annotations.
Main Findings
- Outperforms existing architectures: The abstract states that MambaX-Net significantly outperforms state-of-the-art U-Net and Transformer-based models. The abstract does not report the metrics, effect sizes, or the specific baseline results behind this claim.
- Works with limited and noisy training data: The authors report superior prostate zone segmentation even when the model is trained on limited and noisy data, which is the regime the semi-supervised strategy is designed for. The abstract gives no quantitative evidence for this.
- Evaluated on a longitudinal Active Surveillance dataset: The evaluation was performed on a longitudinal AS dataset, matching the intended clinical use case. The abstract does not state the dataset size, number of time points per patient, or annotation protocol.
- Two components are presented as the source of the gains: The Mamba-enhanced cross-attention module and the shape extractor module are the paper's stated novel components, but the abstract does not provide ablations isolating their individual contributions.
Methodology in Plain English
The model takes two scans rather than one. To segment the current time point, it is given the current MRI alongside the previous MRI and the previous segmentation mask. The idea is that the earlier mask carries useful anatomical information — the prostate's shape and zone boundaries — that constrains and stabilises the new prediction, and the earlier image tells the model how things have changed.
Two modules handle that information. The cross-attention module lets the current scan "look at" the previous one to compare features across time, and a Mamba block is folded into this attention so the model can handle long-range spatial relationships and the temporal comparison without becoming prohibitively expensive. The shape extractor module takes the previous mask and compresses it into a latent representation of anatomy, which is used to sharpen the delineation of prostate zones.
Training is where the label scarcity is addressed. A pre-trained nnU-Net is used to generate pseudo-labels, and these serve as supervision for a self-training procedure, so the model does not require a large set of expert-annotated scans. The abstract does not describe the self-training schedule, confidence thresholds, how many scans were pseudo-labelled versus expert-labelled, or the loss functions used.
Why This Matters
Impact on research. The work pushes segmentation away from the single-time-point assumption that dominates medical imaging benchmarks and toward longitudinal, prior-informed models. It also adds to a growing line of work bringing Mamba/state-space models into medical imaging, and it applies semi-supervised self-training in a setting where expert labels are genuinely the bottleneck rather than a convenience. Both the dual-scan formulation and the pseudo-label pipeline are reusable ideas beyond prostate MRI.
Real-world applications:
- Active Surveillance monitoring: automating the serial prostate measurements that AS protocols depend on, so progression can be tracked consistently across visits.
- Reducing manual contouring burden: radiologists and radiation oncologists currently contour prostate zones by hand; a model that carries anatomy forward from a prior scan can cut that effort.
- Downstream cancer detection and diagnosis: the paper frames segmentation as a preliminary step that enables automated detection and diagnosis of prostate cancer, so improvements here propagate to those tasks.
- Longitudinal change quantification: having comparable segmentations at each time point makes it easier to measure how the gland and its zones change over time.
Industry relevance. Clinical imaging software vendors and radiology AI companies building prostate or oncology workflow tools would be the direct beneficiaries, particularly those targeting surveillance pathways where scan volumes are high, reimbursement pressure is real, and the label budget for bespoke training data is low. The semi-supervised design — bootstrapping from an existing model such as nnU-Net — is also attractive commercially because it lowers the cost of adapting a model to a new site or scanner. No deployment, regulatory, or cost analysis appears in the abstract.
Future Directions
- Quantifying the gains properly: the abstract claims significant improvement over U-Net and Transformer baselines but reports no numbers; full metric tables and component ablations are the obvious next step to establish how much each module contributes.
- Testing generalisation: the abstract describes evaluation on a single longitudinal AS dataset, leaving open whether the dual-scan design transfers to other cohorts, scanners, institutions, or prostate MRI protocols.
- Extending past two time points: the architecture as described conditions on the immediately previous time point; whether it can usefully aggregate many visits, or cope with missed or irregularly spaced scans, is unaddressed.
- Addressing the pseudo-label dependency: the self-training strategy rests on nnU-Net pseudo-labels, so how sensitive the final model is to their quality — and whether the approach can be made less reliant on that particular teacher — is an open question.
- Clinical validation: showing that segmentation quality translates into reliable surveillance decisions, or into downstream detection performance, would require reader studies and outcome-linked evaluation that the abstract does not mention.
Target Audience
Researchers and graduate students working on medical image segmentation, longitudinal or multi-time-point medical imaging, and semi-supervised learning under label scarcity. It is also relevant to engineers building state-space model (Mamba) architectures for 3D or volumetric vision tasks, and to clinical and translational teams working on prostate cancer Active Surveillance who want to understand what automated segmentation can currently offer their workflow. Readers looking for benchmark numbers or deployment evidence will not find them in the abstract.
Authors’ abstract
Active Surveillance (AS) is a treatment option for managing low and intermediate-risk prostate cancer (PCa), aiming to avoid overtreatment while monitoring disease progression through serial MRI and clinical follow-up. Accurate prostate segmentation is an important preliminary step for automating this process, enabling automated detection and diagnosis of PCa. However, existing deep-learning segmentation models are often trained on single-time-point, expertly annotated datasets, making them unsuitable for longitudinal AS analysis, where multiple time points and a scarcity of expert labels hinder effective fine-tuning. To address these challenges, we propose MambaX-Net, a novel semi-supervised, dual-scan 3D segmentation architecture that computes the segmentation for time point t by leveraging the MRI and the corresponding segmentation mask from the previous time point. We introduce two new components: (i) a Mamba-enhanced Cross-Attention Module, which integrates the Mamba block into cross-attention to efficiently capture temporal evolution and long-range spatial dependencies, and (ii) a Shape Extractor Module that encodes the previous segmentation mask into a latent anatomical representation for refined zone delineation. Moreover, we use a semi-supervised self-training strategy that leverages pseudo-labels generated from a pre-trained nnU-Net, enabling effective learning without expert annotations. MambaX-Net was evaluated on a longitudinal AS dataset, and results showed that it significantly outperforms state-of-the-art U-Net and Transformer-based models, achieving superior prostate zone segmentation even when trained on limited and noisy data.