Skip to content
AI.info

Research

ProSona: Prompt-Guided Personalization for Multi-Expert Medical Image Segmentation

Overview Research area: Medical image segmentation (computer vision applied to radiology), specifically multi-rater variability and prompt-guided personalization. Technical level: Intermediate. The pa

ProSona: Prompt-Guided Personalization for Multi-Expert Medical Image Segmentation
arXiv
2511.08046
Published
2025-11-11
Authors
Aya Elgebaly, Nikolaos Delopoulos, Juliane Hörner-Rieber, Carolin Rippke, Sebastian Klüter, Luca Boldrini, Lorenzo Placidi, Riccardo Dal Bello, Nicolaus Andratschke, Michael Baumgartl, Claus Belka, Christopher Kurz, Guillaume Landry, Shadi Albarqouni

AI summary

Overview

Research area: Medical image segmentation (computer vision applied to radiology), specifically multi-rater variability and prompt-guided personalization.

Technical level: Intermediate. The paper assumes familiarity with U-Nets, latent-variable generative segmentation models, and contrastive learning, but the two-stage design and the prompt mechanism are described clearly enough for a reader with basic deep learning background.

Scope (one sentence): ProSona is a two-stage framework that builds a continuous latent space of annotator styles with a probabilistic U-Net and then lets a frozen CLIP text encoder steer that space through natural-language prompts, evaluated on the LIDC–IDRI lung nodule dataset and an in-house multi-institutional prostate MRI dataset.

What This Paper Is About

Medical image boundaries are genuinely ambiguous: radiologists looking at the same lung nodule often draw different contours, and institutions follow different guidelines for prostate MRI. Most automated segmentation methods hide this disagreement by fusing all expert masks into one consensus mask (for example via majority voting), which can erase subtle but clinically meaningful regions such as infiltrative tumor margins that suggest microscopic invasion.

The goal of this paper is to build a segmentation model that keeps expert diversity instead of erasing it, and that a clinician can steer using plain language. The authors want a user to type a prompt such as "conservative segmentation" or "inclusive segmentation" and get a mask matching that style, without needing a separate model branch per annotator.

Key Contributions

  1. A prompt-guided personalization module that maps language descriptions of annotators into a shared latent space and generates similarity-weighted latent codes for personalized predictions.

  2. A multi-level contrastive objective that aligns textual and visual (latent) representations, improving style disentanglement and enabling smooth interpolation between annotators.

  3. Extensive evaluation on the LIDC–IDRI lung nodule dataset and a multi-institutional in-house prostate MRI dataset, showing superior fidelity to individual annotators and styles.

  4. An open implementation released at https://github.com/albarqounilab/ProSona.

Main Findings

  • Lower Generalized Energy Distance on LIDC–IDRI: ProSona reaches a GED of 0.120, versus 0.144 for DPersona, 0.150 for Pionono, 0.232 for TAB, 0.241 for CM-Pixel, and 0.243 for CM-Global. The abstract reports this as a 17% reduction relative to DPersona.

  • Higher Dice scores on LIDC–IDRI: ProSona records Dice Soft 91.56, Dice Max 92.29, Dice Match 90.26, and Mean Dice 90.26, compared with DPersona's 90.31, 90.38, 89.17, and 89.17. The authors describe this as a mean Dice improvement of more than one point.

  • Single-annotator U-Nets trail the multi-rater methods: the four per-annotator U-Nets score GED 0.306 (A1), 0.246 (A2), 0.244 (A3), and 0.296 (A4).

  • Generalization to prostate MRI: on the multi-institutional prostate dataset, ProSona reaches GED 0.146, Dice Soft 87.74, Dice Max 90.26, Dice Match 87.02, and Mean Dice 87.03, versus DPersona's GED 0.158, Dice Soft 87.70, Dice Max 90.01, Dice Match 86.96, and Mean Dice 86.96.

  • Contrastive weighting matters for GED: sweeping alpha and beta over {0, 0.5, 1} shows that increasing the similarity-level weight beta consistently lowers GED on LIDC–IDRI, while prostate MRI performance stays relatively robust across these settings.

  • Prompts reproduce expert tendencies qualitatively: given prompts such as "conservative radiologist" or "inclusive radiologist," ProSona reproduces subtle boundary differences matching expert behavior, and shifting a prompt from "small nodule only" to "include subtle regions" progressively expands the segmentation. The authors note this smooth interpolation is not attainable with discrete multi-head baselines.

Methodology in Plain English

The framework runs in two sequential stages, both trained on the same U-Net-style backbone with a six-dimensional latent code.

Stage 1 — learn the space of annotation styles. The model follows the Probabilistic U-Net design. An encoder produces a deterministic feature map, and two convolutional heads estimate the mean and variance of a Gaussian prior. A separate posterior network, given both the image and a randomly chosen expert annotation, produces a second Gaussian. A latent code sampled from the posterior is concatenated with the encoder features and decoded into a segmentation. Training combines three terms: a Dice segmentation loss against the chosen annotation, a KL divergence pulling the posterior toward the prior, and a boundary loss that pushes the ensemble of sampled predictions to match the experts' intersection and union regions. That boundary term is what prevents the model from collapsing into a single averaged mask.

Stage 2 — let prompts navigate that space. A frozen, pre-trained CLIP text encoder embeds the prompt, and a two-layer MLP projects it into the latent space. The model draws a set of latent samples from the prior and scores each one by dot-product similarity with the projected prompt (scaled by the square root of the latent dimension). Those similarities are turned into softmax weights, and the weighted sum becomes the prompt-specific latent code. This code is concatenated with the deterministic U-Net features and decoded into the personalized mask. Because the combination is a soft weighting rather than a hard selection, the model interpolates smoothly between annotator styles.

Keeping text and latent space consistent. Two contrastive losses operate on batches of annotator prompts. One compares prompt embeddings to each other; the other compares the similarity profiles (each prompt's similarities to the K latent samples) to each other. Both are binary cross-entropy losses over a mask that marks which prompt pairs come from the same annotator, with a temperature parameter. These are added to the segmentation loss, weighted by alpha and beta.

Experimental setup. LIDC–IDRI contributes 1,609 two-dimensional CT slices, each annotated by four radiologists, with annotators ordered from conservative to inclusive by mask area following DPersona. The in-house prostate MRI data comes from four clinical centres (LMU University Hospital, Gemelli University Hospital, University Hospital Zurich, Heidelberg University Hospital). Because each prostate slice was labelled by a single expert, the authors generate three additional pseudo-masks per image using independent nnU-Net models trained on each institution's data, which emulates multi-rater variability while preserving privacy. Images are resampled to isotropic spacing, intensity-normalized, and cropped to 128 × 128, giving a final dataset of 4,817 slices. Four-fold cross-validation is performed at the patient level. Both stages train for 100 epochs with Adam, learning rate 10⁻⁴, batch size eight, and K = 10 latent samples per image.

Why This Matters

Impact on research. The paper reframes multi-rater segmentation from a consensus problem into a controllability problem: instead of predicting one mask or requiring one decoder branch per annotator, it offers a single navigable manifold that can be addressed with language. That connects multi-rater segmentation to vision-language representation learning and gives a concrete alternative to the image-dependent style selection used by DPersona, which the authors criticize as architecturally inefficient with weak style disentanglement.

Real-world applications.

  • Radiotherapy planning, where contouring decisions for prostate and lung targets directly affect dose delivery and where guideline differences between institutions matter.
  • Contour review in clinical workflows, where a reviewer could prompt for a conservative or inclusive interpretation and inspect how the boundary changes.
  • Lung nodule assessment in CT screening, where agreement on nodule extent is known to be poor and where infiltrative extensions can influence aggressiveness estimates.
  • Multi-institutional research studies that need to preserve rather than average away site-specific annotation conventions.

Industry relevance. The method targets a real deployment constraint: per-annotator model branches scale badly as the number of readers, sites, or guidelines grows, whereas a prompt-conditioned model with a shared backbone offers one artifact to maintain. Prompt-based steering also gives a natural interface for regulatory review and model explainability, and the released implementation lowers the barrier to adoption.

Future Directions

  • Richer prompt formulations than simple style descriptions such as "conservative mask" or "inclusive mask," which the authors list explicitly as future work.
  • Extension from 2D slices to 3D volumetric data, also named as future work.
  • Joint modeling of visual and textual uncertainty to support active annotation guidance.
  • Open question raised by the setup: the prostate experiments rely on nnU-Net pseudo-masks to simulate multiple readers, so how the approach behaves on genuinely multi-reader prostate annotations is not established here.

Target Audience

Researchers and practitioners in medical image analysis who work on multi-rater or ambiguous segmentation, including those building radiotherapy and radiology planning tools. It is also relevant to engineers working on vision-language models in clinical settings, and to clinicians interested in how inter-observer disagreement can be represented and steered rather than averaged away. Readers should be comfortable with U-Net architectures, latent-variable models, and contrastive objectives.

Authors’ abstract

Automated medical image segmentation suffers from high inter-observer variability, particularly in tasks such as lung nodule delineation, where experts often disagree. Existing approaches either collapse this variability into a consensus mask or rely on separate model branches for each annotator. We introduce ProSona, a two-stage framework that learns a continuous latent space of annotation styles, enabling controllable personalization via natural language prompts. A probabilistic U-Net backbone captures diverse expert hypotheses, while a prompt-guided projection mechanism navigates this latent space to generate personalized segmentations. A multi-level contrastive objective aligns textual and visual representations, promoting disentangled and interpretable expert styles. Across the LIDC-IDRI lung nodule and multi-institutional prostate MRI datasets, ProSona reduces the Generalized Energy Distance by 17% and improves mean Dice by more than one point compared with DPersona. These results demonstrate that natural-language prompts can provide flexible, accurate, and interpretable control over personalized medical image segmentation. Our implementation is available online 1 .

Read the original paper