Research
Longitudinal Vestibular Schwannoma Dataset with Consensus-based Human-in-the-loop Annotations
Overview Research area: Medical image analysis / computer vision for healthcare, specifically MRI segmentation of vestibular schwannoma (VS) tumours and the construction of a publicly released longitu

- arXiv
- 2511.00472
- Published
- 2025-11-01
- Authors
- Navodini Wijethilake, Marina Ivory, Oscar MacCormac, Siddhant Kumar, Aaron Kujawa, Lorena Garcia-Foncillas Macias, Rebecca Burger, Amanda Hitchings, Suki Thomson, Sinan Barazi, Eleni Maratos, Rupert Obholzer, Dan Jiang, Fiona McClenaghan, Kazumi Chia, Omar Al-Salihi, Nick Thomas, Steve Connor, Tom Vercauteren, Jonathan Shapey
AI summary
Overview
Research area: Medical image analysis / computer vision for healthcare, specifically MRI segmentation of vestibular schwannoma (VS) tumours and the construction of a publicly released longitudinal dataset.
Technical level: Intermediate. The clinical motivation is accessible to beginners, but the paper assumes familiarity with segmentation metrics (Dice Similarity Coefficient), the nnU-Net framework, bootstrapping/fine-tuning, and domain shift concepts.
Scope: The paper describes a consensus-based, human-in-the-loop deep learning pipeline used to create and quality-control a 190-patient longitudinal MRI dataset of vestibular schwannoma, reports the segmentation accuracy achieved across three bootstrapping rounds, and releases the dataset on The Cancer Imaging Archive (TCIA).
What This Paper Is About
Accurately outlining (segmenting) vestibular schwannoma tumours on MRI is needed for clinical management, but manual annotation by experts is slow, expensive, and subjective, while automated deep learning models often fail to generalise from one hospital's imaging protocol to another's. The authors build a pipeline that combines a 3D nnU-Net model with multiple rounds of expert review and correction to efficiently produce trustworthy annotations for a new multi-centre routine clinical dataset, and they measure how much accuracy improves and how much time is saved compared with fully manual annotation. The resulting dataset is released publicly.
Key Contributions
- A human-in-the-loop annotation framework combining a 3D nnU-Net segmentation model, multi-round bootstrapping, consensus-driven expert review, and a post-hoc expert validation stage on diverse datasets.
- A demonstration that bootstrapped fine-tuning on a target internal dataset raises median Dice Similarity Coefficient from 0.9125 to 0.9670 while external test performance stays stable (0.9298 to 0.9229 across rounds).
- A quantified efficiency analysis indicating the proposed approach takes 6.35 hours for 427 cases versus approximately 10.15 hours for manual annotation, a 37.4% reduction.
- Public release of the curated UK MC-RC-2 longitudinal dataset on TCIA, including tumour annotations, demographics, and acquisition metadata.
Main Findings
-
Internal accuracy improved across rounds: On the UK MC-RC-2 internal validation set, median DSC rose from 0.9125 (CI: [0.8075, 0.9483]) in Round 1 to 0.9639 (CI: [0.8931, 0.9844]) in Round 2 and 0.9670 (CI: [0.8671, 0.9915]) in Round 3. A Wilcoxon signed-rank test found the improvement statistically significant (p < 0.05) between Rounds 1 and 2 and between Rounds 1 and 3; the Round 2 to Round 3 improvement was not statistically significant, suggesting a plateau.
-
External performance remained stable: On the external test set (50 scans from UK MC-RC, ETZ SC-GK, and LDN SC-GK), median DSC was 0.9298 (CI: [0.8958, 0.9491]) in Round 1, 0.9265 (CI: [0.8981, 0.9533]) in Round 2, and 0.9229 (CI: [0.8947, 0.9531]) in Round 3, indicating internal bootstrapping did not significantly affect external generalisation.
-
Efficiency gain of 37.4%: The proposed multi-expert pipeline took 6.35 hours for 427 cases versus an estimated 10.15 hours for manual annotation. In Round 1, on 30 randomly selected cases, median quality-assessment time per case was 21.13 seconds (neurosurgical trainee 1), 29.95 seconds (neurosurgical trainee 2), and 25.70 seconds (trained radiologist). In Round 2, over 20 cases, median full manual annotation took 85.65 seconds versus 72.40 seconds for correcting DL segmentations, a difference that was not statistically significant.
-
Annotation sorting across rounds: Of 427 DL-generated proposals in Round 1, 221 were accepted, 156 rejected, 32 excluded, and 18 difficult cases forwarded for expert opinion. Round 2 processing of the 156 rejected sessions yielded 11 accepted, 130 requiring correction, 7 forwarded for expert opinion, and 8 excluded (6 for absent contrast uptake on T1, 2 for cropped tumour regions).
-
Expert review of 143 scans: An expert neuroradiologist reviewed 286 segmentations (143 each from Model 1 and Model 2). Agreement results: 82 accepted by both models, 19 accepted by Model 2 but rejected for Model 1, 16 accepted by Model 1 but rejected for Model 2, and 26 rejected by both. A paired Wilcoxon signed-rank test with Bonferroni correction between the two models yielded p < 0.0001 for DSC, yet expert accept/reject decisions remained consistent across models.
-
Failure modes are qualitative, not just metric-driven: Table 3 documents rejection reasons including over-contouring of small intracanalicular tumours, missed discrete tumour regions, failure to detect extension into the cochlea, vessel mis-segmentation, missed residual tumour in post-operative cases, and cases where both models failed on different slices of the same scan.
-
Dataset composition: The finalised dataset includes 190 patients. The abstract reports tumour annotations for 534 longitudinal T1CE scans from 184 patients plus non-annotated T2-weighted scans from 6 patients; the finalised-dataset text reports 184 patients with 533 annotated T1CE images; the data records section reports 543 T1CE scans, 481 T1-weighted scans, and 133 T2-weighted scans across 621 time points (mean 3.25 scans per patient, mean monitoring period 4.83 ± 3.08 years), with masks provided for 534 T1CE scans and no masks for 9 post-operative scans lacking visible residual tumour.
-
Cohort characteristics: 82 males and 79 females, with sex information unavailable for 29 patients; mean age at first MRI was 58.1 years. Most common symptom was hearing loss (175 patients), followed by tinnitus (56), balance disturbance (48), facial numbness (25), facial weakness (23), headache (15), dizziness (14), and speech difficulties (2). Treatment history: 14 patients had surgery alone, 37 had stereotactic radiosurgery alone, and 2 had both.
Methodology in Plain English
The team started with three already-annotated external datasets (UK MC-RC, a ten-site UK routine clinical collection; ETZ SC-GK, a single-centre gamma knife dataset from Tilburg; and LDN SC-GK, a single-centre gamma knife dataset from London). They trained the default 3D full-resolution nnU-Net on these to produce an initial segmentation model, then ran it on 427 T1CE scans from their new target dataset (UK MC-RC-2).
Three experts (two neurosurgical fellows and a trained radiologist) independently marked each automated proposal as Accept, Reject, or Other. Anything marked Other went to a consensus meeting that included a consultant neurosurgeon, where sessions were reclassified as Accept, Reject, or Exclude.
Accepted scans were added to the training pool and the model was retrained — this is the "bootstrapping" step. Rejected scans were re-processed with the updated model, and a trained radiologist either accepted or manually corrected the new output using ITK-SNAP. This cycle repeated for a third round, adding the corrected Round 2 annotations. A further expert review with a neuroradiologist of 23 years' experience resolved 25 complex scans, after which 14 were corrected, 7 were judged to have no definitive tumour (post-operative cases), and 4 were excluded as not being VS.
Finally, to test generalisation, they compared two models on a randomly selected set of 45 patients (143 scans): Model 1, trained only on the heterogeneous UK MC-RC data, and Model 2, trained on UK MC-RC plus the homogeneous gamma knife datasets plus the remaining UK MC-RC-2 scans. The neuroradiologist reviewed all 286 resulting segmentations in randomised, blinded order.
Ethical approval was granted by the NHS Health Research Authority and Research Ethics Committee (18/LO/0532); because patients were selected retrospectively and images were fully anonymised, informed consent was not required.
Why This Matters
Impact on research. The paper provides a concrete, reproducible protocol for building large trustworthy medical imaging datasets when expert annotation time is the bottleneck, and it provides a public longitudinal VS dataset with consensus-based annotations for others to train and benchmark against. It also delivers a cautionary result: a statistically significant DSC improvement between two models did not translate into a change in expert accept/reject decisions, which challenges reliance on DSC alone.
Real-world applications:
- Automated tumour segmentation in routine VS surveillance MRI, reducing the manual contouring burden on radiologists and clinicians.
- Volumetric tumour growth tracking, which is relevant because a 20% volume increase is the current standard threshold used to define VS growth.
- Treatment planning support for stereotactic radiosurgery, where accurate delineation directly affects the planning workflow.
- Quality assurance and triage, where a model flags cases needing expert attention and experts correct only those cases.
Industry relevance. The pipeline reduces annotation cost per case, which matters to any organisation building medical imaging datasets or products. The result that a model fine-tuned on a target distribution improves locally without degrading external performance supports a practical deployment strategy for clinical software that must handle heterogeneous scanners and protocols. The paper also states that two authors are co-founders and shareholders of Hypervision Surgical, while confirming that no products from that company were used and it has no interest in the work.
Future Directions
- Integrating the bootstrapping framework with real-time human-in-the-loop correction so segmentations can be interactively adapted within clinical workflows, as the authors explicitly suggest for future implementations.
- Developing evaluation metrics beyond DSC that better align with expert qualitative judgement, since expert decisions did not change despite the significant DSC improvement between models.
- Investigating interactive, case-by-case correction for complex anatomy that automated generalisation misses, given the plateau observed between Rounds 2 and Round 3.
- Extending or replicating the approach across other tumour types and imaging protocols, particularly post-operative cases and cases with multiple coexisting pathologies such as VS with meningioma, which the authors identify as confounding factors.
Target Audience
Researchers and engineers working on medical image segmentation and dataset curation; radiologists, neurosurgeons, and clinical teams involved in VS surveillance and radiosurgery planning; and machine learning practitioners interested in human-in-the-loop annotation workflows, domain shift, and the gap between quantitative segmentation metrics and expert clinical judgement.
Authors’ abstract
Accurate segmentation of vestibular schwannoma (VS) on Magnetic Resonance Imaging (MRI) is essential for patient management but often requires time-intensive manual annotations by experts. While recent advances in deep learning (DL) have facilitated automated segmentation, challenges remain in achieving robust performance across diverse datasets and complex clinical cases. We present an annotated dataset stemming from a bootstrapped DL-based framework for iterative segmentation and quality refinement of VS in MRI. We combine data from multiple centres and rely on expert consensus for trustworthiness of the annotations. We show that our approach enables effective and resource-efficient generalisation of automated segmentation models to a target data distribution. The framework achieved a significant improvement in segmentation accuracy with a Dice Similarity Coefficient (DSC) increase from 0.9125 to 0.9670 on our target internal validation dataset, while maintaining stable performance on representative external datasets. Expert evaluation on 143 scans further highlighted areas for model refinement, revealing nuanced cases where segmentation required expert intervention. The proposed approach is estimated to enhance efficiency by approximately 37.4% compared to the conventional manual annotation process. Overall, our human-in-the-loop model training approach achieved high segmentation accuracy, highlighting its potential as a clinically adaptable and generalisable strategy for automated VS segmentation in diverse clinical settings. The dataset includes 190 patients, with tumour annotations available for 534 longitudinal contrast-enhanced T1-weighted (T1CE) scans from 184 patients, and non-annotated T2-weighted scans from 6 patients. This dataset is publicly accessible on The Cancer Imaging Archive (TCIA) (https://doi.org/10.7937/bq0z-xa62).