Research
Adaptable Segmentation Pipeline for Diverse Brain Tumors with Radiomic-Guided Subtyping and Lesion-Wise Model Ensemble
Overview Research area: Medical image analysis / computer vision — automatic brain tumor segmentation from multi-parametric MRI, applied to the four tasks of the BraTS 2025 Lighthouse Challenge. Techn

- arXiv
- 2512.14648
- Published
- 2025-12-16
- Authors
- Daniel Capellán-Martín, Abhijeet Parida, Zhifan Jiang, Nishad Kulkarni, Krithika Iyer, Austin Tapp, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru
AI summary
Overview
Research area: Medical image analysis / computer vision — automatic brain tumor segmentation from multi-parametric MRI, applied to the four tasks of the BraTS 2025 Lighthouse Challenge.
Technical level: Intermediate to Advanced. The paper assumes familiarity with deep learning segmentation backbones (nnU-Net, MedNeXt), cross-validation, ensembling, and radiomics, though its pipeline logic can be followed by a general reader.
Scope (one sentence): The authors describe a modular, backbone-independent pipeline — radiomic-guided fold splitting, lesion-wise model selection and weighted ensembling, and cluster- and label-specific post-processing — that they applied unchanged to four different brain tumor cohorts (pediatric tumors, preoperative meningioma, meningioma radiotherapy, and metastases).
What This Paper Is About
Automated segmentation of brain tumors on MRI is hard to make robust because tumor types look very different from one another and annotations vary across scanners and populations. The authors argue that state-of-the-art segmentation networks have converged to very similar Dice scores — a performance plateau — so further gains must come from pipeline engineering rather than from ever more complex architectures. Their goal is a single adaptable pipeline, not tied to any one network, that tailors training, model weighting, and post-processing to each tumor subtype and lesion.
Key Contributions
- A backbone-independent, adaptable four-stage pipeline comprising (i) training data preparation via stratified fold splitting from radiomic clustering, (ii) model training, (iii) model selection and ensembling, and (iv) adaptive post-processing (small connected-component removal plus label redefinition). The same pipeline was applied to all four BraTS 2025 Lighthouse tasks.
- Radiomic-guided stratified cross-validation. For each case, 14 shape-based and 93 appearance-based radiomic features were computed per MRI sequence with PyRadiomics using the whole-tumor mask; principal components explaining 90% of the variance were retained, cases were partitioned by k-means clustering (number of clusters chosen by maximizing the silhouette coefficient), and each cluster was randomly split into five folds.
- A single lesion-wise ranking objective F that aggregates Dice, normalized surface distance (NSD) and Hausdorff distance across tumor sub-regions by case-wise ranking, matching the official BraTS leaderboard strategy. It is used both to weight models in the ensemble and to tune post-processing thresholds. The ranker is open-sourced.
- Adaptive, cluster- and label-specific post-processing: a grid search over connected-component volumes from 0 to 500 voxels in 25-voxel increments, and a volume-ratio-based label-redefinition step derived from a confusion matrix of frequently swapped labels.
Main Findings
- Weighted ensemble outperforms single backbones. Each model's ensemble weight is computed as the sum of the other models' ranking scores divided by the sum of all ranking scores. Weights were nearly even for PED (0.332 nnU-Net V2, 0.331 nnU-Net ResEnc M, 0.337 MedNeXt M) and MEN-RT (0.336, 0.338, 0.326), and more skewed for MET, where MedNeXt M received 0.409 versus 0.295 and 0.296.
- Post-processing produced the largest single-metric gains. For PED validation, the ensemble reached lesion-wise Dice of 0.677 for enhancing tumor and 0.933 for whole tumor, while the post-processed output reached 0.967 NSD for edema (versus 0.868 for the ensemble). On MET validation, post-processing raised resection-cavity Dice from 0.705 to 0.907 and NSD from 0.668 to 0.871.
- Validation-set gains can be modest while cross-validation gains are larger. The authors report that some improvements in Dice or NSD were limited on the validation set but became evident under 5-fold CV, which they attribute to the validation set being smaller than the training set and potentially skewed in tumor-type distribution.
- Selected validation results across tasks: PED ensemble lesion-wise Dice for whole tumor 0.933 and tumor core 0.933; MEN post-processing Dice of 0.856, 0.855 and 0.848 with NSD 0.873, 0.865 and 0.854 for the reported regions; MEN-RT ensemble Dice 0.796 and NSD 0.654, unchanged by post-processing.
- Testing-set means: PED testing mean lesion-wise Dice ranged from 0.669 (cystic component) to 0.915 (whole tumor); MET testing means of 0.544 for enhancing tumor, 0.555 for tumor core, 0.561 for whole tumor and 0.860 for resection cavity; MEN testing means of 0.881, 0.878 and 0.868 with NSD 0.885, 0.878 and 0.860; MEN-RT testing mean Dice 0.807 and NSD 0.683.
- STAPLE fusion behaved differently by task. It lagged behind the weighted ensemble when the three 5-fold-averaged predictions were used on the training data, but performed better on the MEN-RT validation set when each fold's model was treated as an independent candidate, yielding 15 predictions, which markedly improved DSC, NSD and rank for that cohort.
- Qualitative examples: median lesion-wise or global whole-tumor Dice of 0.979 (PED, lesion-wise), 0.868 (MET, global), 0.961 (MEN, lesion-wise) and 0.887 (MEN-RT, lesion-wise).
- Evaluation conditions: validation-set evaluation was performed on the Synapse platform with no access to reference standard annotations on the validation set and no access to any testing data, including images and labels.
Methodology in Plain English
The team trained three well-established segmentation networks — nnU-Net V2, nnU-Net with a residual encoder (nnU-Net ResEnc M), and MedNeXt M (k=3, 17.6M parameters, 248 GFlops) — on each of the four BraTS 2025 datasets using 5-fold cross-validation, with input patches of 128x128x128 voxels. All models used label-wise softmax activation, a class-weighted loss combining Dice and cross-entropy, stochastic gradient descent with Nesterov momentum (initial learning rate 0.01, momentum 0.99, weight decay 3e-5), 200 epochs, and NVIDIA A100 (40 GB) GPUs.
Instead of splitting cases into folds randomly, they first characterized each training case by radiomic features computed from its whole-tumor mask, reduced those features with principal component analysis, clustered the cases with k-means, and then split each cluster into five folds. This is meant to keep tumor appearance diversity balanced across folds.
Because a single metric like Dice is not enough — the challenge scores use Dice, normalized surface distance and Hausdorff distance across several tumor sub-regions — they borrowed BraTS's ranking approach: compute lesion-wise metrics for every candidate prediction, rank predictions case by case, and average these ranks. This produces one internal score F per model (lower is better). Models with better F receive higher weights in a weighted average of their 5-fold probability outputs.
Finally, they re-ran the radiomic clustering on predicted masks so that post-processing thresholds could differ per cluster and per label. The first post-processing step removes small isolated connected components using a grid search over volumes of 0 to 500 voxels in 25-voxel steps, selecting the threshold with the best average rank. The second builds a confusion matrix of frequently swapped labels and, for each confused label pair within each cluster, searches for a threshold on the ratio of the smaller label's volume to whole-tumor volume; predictions below the threshold have that label converted to the other. At inference, a new test case is assigned to its nearest radiomic cluster and the corresponding thresholds are applied.
Why This Matters
Impact on research. The work argues that when segmentation architectures have plateaued, the remaining gains lie in target-specific pipeline design — heterogeneity-aware fold splitting, lesion-aware ensemble weighting, and subtype-specific post-processing — rather than in deeper networks. It also makes a methodological point about evaluation: tuning models and thresholds on a small validation set can be misleading, and cross-validation on the larger training set plus a single aggregated ranking metric is more robust. Releasing the ranker and keeping all components model-independent supports reuse.
Real-world applications:
- Quantitative tumor measurement in neuro-oncology workflows, supporting diagnosis, treatment planning, and longitudinal monitoring of tumor growth.
- Radiotherapy planning for meningioma, where gross tumor volume delineation (the MEN-RT task) directly drives treatment.
- Pre- and post-treatment assessment of brain metastases, including delineation of the resection cavity in post-treatment cases.
- Pediatric neuro-oncology, where the PED task covers multi-consortium international pediatric brain tumor data and includes cystic components alongside enhancing and non-enhancing tumor.
Industry relevance. The pipeline is distributed as ready-to-use Docker containers and a web application, which lowers the barrier to deployment for clinical and commercial imaging platforms. Because the approach is not locked to a specific network architecture, vendors and hospital systems can swap in their own backbones while retaining the selection, ensembling, and post-processing logic.
Future Directions
- Replacing the simple volume thresholds used in both post-processing steps with radiomic-driven criteria analogous to those used for stratification, so post-processing is tailored to tumor morphology and appearance rather than volume alone.
- Running additional ablations across multiple BraTS tasks to quantify how much heterogeneity-aware sampling (radiomic-cluster fold splitting versus random splits) actually reduces overfitting.
- Exploring alternative ensemble-weighting schemes, since the rank-based score can produce very close values when there are only three candidate models, limiting its sensitivity.
- Validating the pipeline in larger clinical studies to assess whether its potential for clinical impact extends beyond BraTS benchmarks.
Target Audience
Researchers and engineers working on medical image segmentation, particularly those participating in or adapting BraTS-style challenges; medical imaging scientists interested in model ensembling, radiomics, and post-processing heuristics; and clinical translation teams — including those in radiology, radiation oncology, and neuro-oncology — looking for a deployable, architecture-agnostic tumor segmentation pipeline.
Authors’ abstract
Robust and generalizable segmentation of brain tumors on multi-parametric magnetic resonance imaging (MRI) remains difficult because tumor types differ widely. The BraTS 2025 Lighthouse Challenge benchmarks segmentation methods on diverse high-quality datasets of adult and pediatric tumors: multi-consortium international pediatric brain tumor segmentation (PED), preoperative meningioma tumor segmentation (MEN), meningioma radiotherapy segmentation (MEN-RT), and segmentation of pre- and post-treatment brain metastases (MET). We present a flexible, modular, and adaptable pipeline that improves segmentation performance by selecting and combining state-of-the-art models and applying tumor- and lesion-specific processing before and after training. Radiomic features extracted from MRI help detect tumor subtype, ensuring a more balanced training. Custom lesion-level performance metrics determine the influence of each model in the ensemble and optimize post-processing that further refines the predictions, enabling the workflow to tailor every step to each case. On the BraTS testing sets, our pipeline achieved performance comparable to top-ranked algorithms across multiple challenges. These findings confirm that custom lesion-aware processing and model selection yield robust segmentations yet without locking the method to a specific network architecture. Our method has the potential for quantitative tumor measurement in clinical practice, supporting diagnosis and prognosis.