Skip to content
AI.info

Research

A Unified Framework for Joint Detection of Lacunes and Enlarged Perivascular Spaces

Overview Research area: Medical image analysis / computer vision for neuroimaging — specifically multi-task 3D deep learning for cerebral small vessel disease (CSVD) markers. Technical level: Advanced

A Unified Framework for Joint Detection of Lacunes and Enlarged Perivascular Spaces
arXiv
2603.04243
Published
2026-03-04
Authors
Lucas He, Krinos Li, Hanyuan Zhang, Runlong He, Silvia Ingala, Luigi Lorenzini, Marleen de Bruijne, Frederik Barkhof, Rhodri Davies, Carole Sudre

AI summary

Overview

Research area: Medical image analysis / computer vision for neuroimaging — specifically multi-task 3D deep learning for cerebral small vessel disease (CSVD) markers.

Technical level: Advanced. The paper assumes familiarity with 3D U-Net architectures, cross-attention gating, deep supervision, Tversky and Dice loss variants, and distance-field-based post-processing.

Scope: The paper presents a single multi-task 3D framework that jointly detects enlarged perivascular spaces (EPVS) and lacunes from T1, T2, and FLAIR MRI, trained on the VALDO 2021 challenge dataset (N=40) and evaluated externally on the EPAD cohort (N=1762).

What This Paper Is About

Lacunes and enlarged perivascular spaces are two different biomarkers of cerebral small vessel disease, but they look confusingly similar on MRI: both appear isointense to cerebrospinal fluid, so software has trouble telling the oval lacunes apart from the tubular EPVS. Most existing tools tackle only one of the two markers at a time, which throws away the fact that the two lesions often occur together, and they struggle with the extreme imbalance between tiny lesions and the rest of the brain.

This paper builds one model that handles both targets at once, using the dense EPVS signal as spatial context to help find the much rarer lacunes, while adding anatomical and topological constraints so the model stops flagging lesions in places where they biologically should not appear.

Key Contributions

  1. Morphology-decoupled architecture: A shared-encoder, dual-decoder 3D network (built on a Dynamic U-Net) with a Zero-Initialized Gated Cross-Task Attention module that sends information one way only — from the EPVS branch to the lacune branch — at every up-sampling stage, so tubular and ovoid features are separated rather than entangled.

  2. Mixed-supervision and anatomy-aware objectives: A hybrid training strategy that combines fully supervised dense masks with weakly supervised regional counts for EPVS, plus a Mutual Exclusion loss to penalize voxel-wise probability overlap between the two classes and a Soft-Centerline Dice loss to preserve the tubular topology of vessels, all balanced with homoscedastic uncertainty weighting.

  3. Anatomically-Informed Inference Calibration: A distance-field-based decision boundary that replaces the standard static 0.5 threshold, raising the confidence required for a voxel to be labelled as lesion the further it sits from anatomically plausible tissue.

  4. Extensive validation: 5-fold cross-validation on VALDO 2021 (N=40) against MONAI baselines (DYN Unet, MedNeXt, Swin-UNETR V2, VISTA-3D) and re-implemented VALDO 2021 task winners, plus external generalization testing on the EPAD cohort (N=1762).

Main Findings

  • Lacune detection beats the challenge winners on precision and F1: The method reaches 71.1 ± 17.3% precision (p = 0.01) and 62.6 ± 17.1% F1-score (p = 0.03) for lacunes on VALDO, both statistically significant against the challenge winner.
  • Lacune segmentation quality: DSC of 42.4 ± 11.3% and NSD of 58.5 ± 15.9%, with false positives reduced to 0.7 ± 0.9 per subject.
  • EPVS detection F1 of 53.7 ± 9.6%, numerically above the VALDO winner (50.5 ± 9.2%), though EPVS recall is lower (49.8 ± 10.0% versus 53.2 ± 4.4%) and EPVS segmentation DSC (38.1 ± 6.5%) and NSD (56.7 ± 8.4%) trail the winner (42.8 ± 9.9% and 63.7 ± 11.5%).
  • External EPAD validation: For lacunes, Balanced Accuracy of 64.9 ± 2.1%, Mean Absolute Error of 0.2 ± 0.01, and global correlation r = 0.24 ± 0.05. For EPVS, Spearman correlations of ρ = 0.22 ± 0.02 (basal ganglia), ρ = 0.29 ± 0.03 (centrum semi-ovale), and ρ = 0.11 ± 0.02 (mid-brain) — the CSO figure substantially exceeds the best baseline, VISTA-3D at ρ = 0.17 ± 0.02.
  • Mixed supervision is essential: Fully supervised single-task learning collapses on EPVS, with an F1 of 18.3 ± 14.8% and 94.8 ± 91.9 false positives per subject; adding mixed supervision recovers F1 to 48.6 ± 9.8%.
  • Gated attention resolves feature interference: Moving from a shared decoder to the gated-attention split improves EPVS F1 from 50.0% to 51.8% and lacune F1 from 56.1% to 58.2%.
  • Each loss term targets a different failure mode: The Mutual Exclusion loss lifts lacune recall from 50.0 ± 12.8% to 68.1 ± 17.5% while cutting false positives from 3.9 to 2.2 per subject; the Soft-Centerline Dice loss pushes EPVS precision to its peak of 62.6 ± 19.1%; the combination yields the lowest lacune false-positive rate at 1.0 ± 0.7.
  • Anatomical calibration trades sensitivity for specificity: Applying inference calibration reduces lacune false positives to 0.7 ± 0.9 with minimal recall change, but the authors note it occasionally misses small or faint lesions.

Methodology in Plain English

The team started from a 3D U-Net style network but split it into two branches after a shared encoder: one branch learns the long, tube-like shapes of EPVS, the other learns the round, oval shapes of lacunes. The key trick is a gate that lets the EPVS branch inform the lacune branch but never the reverse. The gate is initialized to zero, so early in training it behaves as if it is not there, which keeps training stable — the network only starts using the cross-task signal once it has learned something useful.

Because EPVS labels are scarce, the team trained the EPVS branch with a mix of two supervision types: for 12 subjects there were full dense voxel masks, and for 28 subjects only regional counts. The lacunes were fully segmented. Several loss terms push the model toward biologically sensible behaviour: one discourages the two lesion types from predicting probability mass in the same voxel, and another preserves the thin connected structure of vessels.

For post-processing, the team used FastSurfer to divide each brain into three reliability zones — an allowed zone (white matter, deep grey matter, brainstem), a transition zone (hippocampus and cerebellar white matter), and an exclusion zone (cortex, ventricles, extra-cerebral tissue). They then computed a distance map from the allowed zone and made the binarization threshold rise with distance from it, so a prediction deep in the exclusion zone needs a probability above roughly 0.9 to survive. Finally, connected component analysis converts voxel-level predictions into discrete lesion objects.

Training used patch size 96³, AdamW at a learning rate of 1×10⁻⁴ and weight decay of 1×10⁻⁵ for 200 epochs on a single NVIDIA A100 (80GB), with inference via a sliding window of ROI 128³ and 0.6 overlap.

Why This Matters

Impact on research: The work shows that jointly modelling two radiological mimics through a deliberately asymmetric information channel outperforms treating them as isolated problems. It also demonstrates that weak, region-level labels for EPVS can substitute for expensive dense annotation, and that a purely inference-time anatomical prior can cut false positives without retraining — a cheap recipe that other lesion-detection pipelines could adopt.

Real-world applications:

  • Population-scale epidemiology: automatically extracting lacune burden and EPVS visual ratings from thousands of scans, as demonstrated on the 1762-subject EPAD cohort.
  • Clinical trial enrichment: using reliable lesion counts to select or stratify participants with measurable small vessel disease burden.
  • Dementia and stroke risk assessment: lacunes and EPVS are established biomarkers of vascular contribution to cognitive decline.
  • Radiology workflow triage: flagging scans with high lesion burden for prioritised expert review, while suppressing implausible cortical false positives that currently waste reader time.

Industry relevance: Medical imaging software vendors and CROs need automated CSVD quantification that generalizes across scanners and cohorts without per-site retraining. This framework's combination of multi-task efficiency (one shared encoder), tolerance for weak labels, and inference-time calibration maps directly onto that requirement. The authors state the code will be released upon acceptance.

Future Directions

  • Recovering lost sensitivity: The anatomical calibration improves precision but occasionally misses small or faint lesions; finding an adaptive calibration that preserves recall is an explicit open problem.
  • Closing the segmentation gap on EPVS: The method's EPVS DSC (38.1 ± 6.5%) and NSD (56.7 ± 8.4%) remain below the VALDO winner, so better tubular structure recovery is needed.
  • Scaling beyond the VALDO cohort: The training set is only 40 subjects; whether the gains hold with larger and more demographically diverse training data is not established in this paper.
  • Reducing dependence on external anatomical segmentation: The pipeline relies on FastSurfer for zoning; alternatives that learn the anatomical prior end-to-end are a natural extension. The paper does not report inference cost or runtime, which would matter for deployment.

Target Audience

Researchers and engineers working on medical image segmentation and multi-task learning, particularly those handling small, imbalanced, mimic-prone lesions in brain MRI. It is also relevant to neuroimaging scientists studying cerebral small vessel disease, clinical trial methodologists needing automated CSVD biomarkers, and machine learning practitioners interested in weak supervision, cross-task attention gating, and anatomy-informed post-processing. Readers without a background in 3D segmentation architectures or neuroanatomy will find the methodological sections challenging.

Authors’ abstract

Cerebral small vessel disease (CSVD) markers, specifically enlarged perivascular spaces (EPVS) and lacunae, present a unique challenge in medical image analysis due to their radiological mimicry. Standard segmentation networks struggle with feature interference and extreme class imbalance when handling these divergent targets simultaneously. To address these issues, we propose a morphology-decoupled framework where Zero-Initialized Gated Cross-Task Attention exploits dense EPVS context to guide sparse lacune detection. Furthermore, biological and topological consistency are enforced via a mixed-supervision strategy integrating Mutual Exclusion and Centerline Dice losses. Finally, we introduce an Anatomically-Informed Inference Calibration mechanism to dynamically suppress false positives based on tissue semantics. Extensive 5-folds cross-validation on the VALDO 2021 dataset (N=40) demonstrates state-of-the-art performance, notably surpassing task winners in lacunae detection precision (71.1%, p=0.01) and F1-score (62.6%, p=0.03). Furthermore, evaluation on the external EPAD cohort (N=1762) confirms the model's robustness for large-scale population studies. Code will be released upon acceptance.

Read the original paper