Research
Demo: Generative AI helps Radiotherapy Planning with User Preference
Overview Research area: Medical computer vision / generative AI applied to radiotherapy (RT) treatment planning, submitted to the Second Workshop on GenAI for Health: Potential, Trust, and Policy Comp
- arXiv
- 2512.08996
- Published
- 2025-12-08
- Authors
- Riqiang Gao, Simon Arberet, Martin Kraus, Han Liu, Wilko FAR Verbakel, Dorin Comaniciu, Florin-Cristian Ghesu, Ali Kamen
AI summary
Overview
- Research area: Medical computer vision / generative AI applied to radiotherapy (RT) treatment planning, submitted to the Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance (arXiv:2512.08996v1 [cs.CV], 08 Dec 2025).
- Technical level: Intermediate. The clinical motivation and slider concept are accessible, but the method section assumes familiarity with VQ-VAE, GANs, adaptive instance normalization, and dose-volume histograms.
- Scope: The paper presents a two-stage generative model ("flexible dose proposer", FDP) that predicts 3D radiotherapy dose distributions conditioned on user-adjustable preference sliders, and demonstrates integration with a commercial treatment planning system (Eclipse).
What This Paper Is About
Most deep learning dose-prediction models are trained against reference plans, so they absorb the planning style or institutional preferences baked into that ground truth, and a single trained model gives the user no way to steer the trade-off between target coverage and organ-at-risk sparing. This paper builds a generative model that predicts 3D dose from anatomy plus user-defined preference "flavors" — set through interactive sliders — and then converts the predicted dose into optimization objectives that a clinical planning system can execute into a deliverable plan.
Key Contributions
- A two-stage training framework in which a foundational dose decoder is pre-trained in Stage I and reused in Stage II, improving training stability under complex conditions such as user-preference conditioning.
- What the authors describe as the first dose prediction model with interactive sliders, enabling real-time customization of the trade-off between target homogeneity and organ-at-risk (OAR) sparing.
- Integration of the AI model with a widely used clinical treatment planning system, with demonstrated generation of high-quality radiotherapy plans from the predicted dose.
- Comparative evaluation against the Varian RapidPlan model, which the authors state their method surpasses in both adaptability and plan quality in some scenarios.
Main Findings
- Intra-patient DVH differences favor the proposed model. In Table 2, the standard-deviation "better count" is 0 for RapidPlan versus 15 for FDP, and the mean "better count" is 1 versus 14. Individual structures follow the same pattern, for example SpinalCord05 standard deviation of 3.02 for RapidPlan versus 1.18 for FDP, and mean of 3.25 versus 1.69.
- Inter-patient DVH differences also favor FDP. In Table 3, the standard-deviation better count is 3 for RapidPlan versus 12 for FDP, and the mean better count is 2 versus 12.
- Achieved plan quality favors FDP on OAR sparing. Table 4 reports an OAR "better" count of 14 for FDP versus 0 for RapidPlan, and a PTV homogeneity/conformity "better" count of 1 versus 0. The "similar" thresholds used are 1 Gray for OARs and 0.015 for PTV indices.
- Stage I pre-training helps. Test-set MAE masked by 5 Gy isodose lines is 2.63 without Stage I pre-training versus 2.56 with it (Table 6), and the authors report that without Stage I the predicted dose shows boundary artifacts of PTVs and OARs.
- Sliders produce measurably different plans. Table 5 compares two preference settings: Preference 1 (P1, OAR sparing over PTV homogeneity) and Preference 2 (P2, the opposite). For example, SpinalCord05 mean dose is 14.59 under P1 versus 16.33 under P2, and ParotidCon-PTV maximum dose is 12.5 under P1 versus 15.5 under P2.
- Results hold under a second planning technique. The main tables reflect VMAT-based planning; the appendix reports IMRT outcomes with the same direction (Table 11: OAR "better" count 10 versus 4 in favor of FDP, PTV count 3 versus 0).
- RapidPlan's distribution is reflected more on a small validation set. Appendix 10 reports that on small validation sets, RapidPlan's distribution is captured more closely than FDP's, with FDP still leading on most counts (Table 12: standard-deviation better count 9 versus 6; Table 13: 7 versus 5).
- Inference is fast. In the demo, the model forward pass takes about 30 ms, visualization takes about 1.5 seconds, and DICOM loading and pre-processing takes about half a minute, run once per patient.
- Authors weight standard deviation over mean. Because mean differences can be offset by consistent margins with rule-based methods, the paper de-emphasizes mean values and treats large standard deviations as the indicator of poor estimation consistency and reliability.
Methodology in Plain English
The model is trained in two steps. In Stage I, a VQ-VAE learns a compressed latent representation of realistic dose distributions from a large corpus of 31K doses (not necessarily from high-quality plans). The purpose of this stage is not compression for speed but regularization: it gives Stage II a decoder that already produces realistic-looking dose, which stabilizes later training.
In Stage II, the CT image and radiotherapy structures are concatenated as a multi-channel input and encoded in a MedNext-style encoder, while user preferences and beam/angle information are injected through adaptive instance normalization. The model is trained with a combination of image-space reconstruction, latent-space reconstruction, adversarial loss, and an "objective" loss that penalizes inconsistency between the user's slider values and the predicted plan metrics. The objective loss has three parts: matching a target homogeneity index, aligning PTV mean dose, and honoring an OAR-sparing preference weight. During training the slider values are sampled randomly within predefined ranges so the model learns to respond to the full range of inputs.
The authors deliberately avoid diffusion-based generation in favor of one-step GAN generation, which keeps inference fast enough for interactive use. After prediction, the 3D dose is converted into dose-volume and mean-dose objectives that Eclipse can optimize on, since Eclipse does not accept a 3D dose distribution as a direct objective. OAR DVHs are sampled at specific volume percentages and given rule-based margins; PTV and PTV ring objectives follow RapidPlan conventions and come from prescriptions rather than model estimates.
Evaluation uses six cohorts of head-and-neck cancer cases with train/validation/test splits of 370/48/54, 147/17/19, 128/15/17, 103/14/12, 52/8/7, and 20/1/4. The primary metric is the difference between expected and achieved DVHs, computed both within patients and across patients, with lower values and lower standard deviations indicating better estimation. The baseline is a high-quality RapidPlan setup, and all test plans followed RapidPlan structure requirements.
Why This Matters
- Research impact: It reframes dose prediction from a single-answer regression problem into a controllable generation problem, where the model is conditioned on explicit user preferences rather than an implicit institutional style. It also provides a validation pathway that goes beyond prediction accuracy to actual deliverable plans inside a commercial planning system.
- Real-world applications:
- Interactive planning, where a dosimetrist or physicist moves sliders to bias OAR sparing versus target homogeneity and sees an updated dose prediction almost instantly.
- Knowledge-based planning replacement or augmentation, giving institutions flexibility that would otherwise require training their own institution-specific RapidPlan model.
- Quality assurance and plan review, where predicted versus achieved DVHs serve as a check on planning consistency.
- Automation of head-and-neck planning, described by the authors as the most challenging region for RT planning due to complex anatomy and stringent dose constraints.
- Industry relevance: The work comes from Siemens Healthineers (Digital Technology and Innovation) with a co-author from Varian Medical Systems, and the evaluation is performed against Varian RapidPlan and inside the Eclipse treatment planning system, indicating a path toward product-level deployment. The paper carries an explicit disclaimer that the results are not commercially available and that future commercial availability cannot be guaranteed.
Future Directions
- Extending the approach beyond head-and-neck cancer to other treatment sites, which the authors expect to be straightforward and list as future work.
- Conducting more rigorous clinical quality comparisons across diverse clinical scenarios, since the training paradigms of RapidPlan (conventional methods, smaller datasets) and FDP (deep learning, larger-scale data) differ fundamentally and plan quality comparison is inherently complex and subjective.
- Releasing full dataset details, which the authors state will be disclosed after the double-blind review.
- Addressing the observation that RapidPlan's distribution is captured more closely on small validation sets, which raises the question of how FDP behaves when only limited institutional data is available.
Target Audience
Medical physicists, dosimetrists, and radiation oncology researchers interested in AI-assisted treatment planning; machine learning researchers working on conditional generative models for medical imaging; and industry engineers evaluating integration of AI dose prediction into commercial treatment planning systems. Readers need some familiarity with radiotherapy planning concepts (DVHs, PTVs, OARs, HI/CI) and with deep generative architectures to get full value from the method and appendix sections.
Authors’ abstract
Radiotherapy planning is a highly complex process that often varies significantly across institutions and individual planners. Most existing deep learning approaches for 3D dose prediction rely on reference plans as ground truth during training, which can inadvertently bias models toward specific planning styles or institutional preferences. In this study, we introduce a novel generative model that predicts 3D dose distributions based solely on user-defined preference flavors. These customizable preferences enable planners to prioritize specific trade-offs between organs-at-risk (OARs) and planning target volumes (PTVs), offering greater flexibility and personalization. Designed for seamless integration with clinical treatment planning systems, our approach assists users in generating high-quality plans efficiently. Comparative evaluations demonstrate that our method can surpasses the Varian RapidPlan model in both adaptability and plan quality in some scenarios.