Research
Parallel Swin Transformer-Enhanced 3D MRI-to-CT Synthesis for MRI-Only Radiotherapy Planning
Overview Research area: Medical image synthesis and computer vision, specifically MRI-to-CT translation for radiotherapy planning. Technical level: Advanced. The work combines 3D convolutional encodin

- arXiv
- 2602.05387
- Published
- 2026-02-05
- Authors
- Zolnamar Dorjsembe, Hung-Yi Chen, Furen Xiao, Hsing-Kuo Pao
AI summary
Overview
Research area: Medical image synthesis and computer vision, specifically MRI-to-CT translation for radiotherapy planning.
Technical level: Advanced. The work combines 3D convolutional encoding, Swin Transformer attention mechanisms, and dosimetric evaluation, and assumes familiarity with generative image synthesis and radiation oncology workflows.
Scope: The paper proposes a 3D transformer-based architecture for generating synthetic CT volumes from MRI, evaluated on both public and clinical data with image-similarity, geometric, and dose-accuracy measures.
What This Paper Is About
Radiotherapy dose calculation requires electron density information, which MRI does not provide directly, so clinical workflows typically acquire both MRI and CT scans for the same patient. Combining these acquisitions introduces registration uncertainty and extra procedural steps. The paper's goal is to synthesize CT-like images from MRI alone using a 3D deep learning architecture, so that treatment planning can be performed without a separate CT scan.
Key Contributions
- A parallel Swin Transformer-enhanced architecture (Med2Transformer): The model combines convolutional encoding with two parallel Swin Transformer branches, designed to capture both fine local anatomical detail and long-range spatial context in 3D volumes.
- Multi-scale shifted window attention with hierarchical feature aggregation: These mechanisms are introduced to improve anatomical fidelity in the synthesized images.
- Evaluation across public and clinical datasets: The authors report image similarity and geometric accuracy improvements relative to baseline methods.
- Dosimetric validation: The synthesized images are assessed at the dose level, with a reported mean target dose error of 1.69%, which the authors describe as clinically acceptable. Code is released publicly at the linked repository.
Main Findings
- Image similarity: The proposed method achieves higher image similarity than the baseline methods it was compared against, according to the abstract.
- Geometric accuracy: The authors report improved geometric accuracy relative to baselines.
- Dosimetric performance: Dosimetric evaluation indicates clinically acceptable results, with a mean target dose error of 1.69%. This is the only quantitative figure given in the abstract.
- Dataset coverage: Experiments span both public and clinical datasets, though the abstract does not specify their size, composition, or the identity of the baseline methods.
- Availability: The implementation is released as open-source code.
The abstract does not report specific similarity metric values, geometric error figures, baseline names, dataset sizes, or ablation results; those details would require the full text.
Methodology in Plain English
The researchers built a 3D neural network that takes an MRI volume as input and produces a corresponding CT volume. The network splits the work: a convolutional part handles the local, fine-grained anatomy, while two parallel transformer branches learn broader spatial relationships across the whole volume. Transformers are useful here because MRI-to-CT conversion is not a simple linear mapping — the relationship between tissue appearance in MRI and the electron density values needed for dose calculation varies with anatomy and patient. The attention mechanism is applied at multiple scales using shifted windows, and features from different levels of the network are combined hierarchically, which the authors argue helps preserve anatomical structure. The resulting synthetic CTs were then compared to real CTs both in terms of image similarity and geometry, and finally fed into dose calculation to check whether the dose distribution was close enough to what a real CT would produce.
Why This Matters
Impact on research: The work contributes to the broader effort to make MRI-only radiotherapy planning viable, an active area where the central obstacle is the absence of electron density information in MRI. By combining convolutional and transformer components in 3D, it offers a design pattern for other cross-modality medical image synthesis tasks where both local texture and global context matter.
Real-world applications:
- Radiotherapy treatment planning that relies on a single MRI session rather than paired MRI and CT acquisitions.
- Reducing registration uncertainty and patient positioning errors that arise when two separate scans must be aligned.
- Eliminating the ionizing radiation dose associated with planning CT scans, which matters for patients undergoing repeated imaging.
- Streamlining clinical workflow and throughput in radiology and radiation oncology departments by removing a scan from the planning pipeline.
Industry relevance: Medical imaging device manufacturers, radiotherapy planning software vendors, and hospital systems all have a stake in whether MRI-only workflows become standard, since it affects both equipment utilization and planning software requirements. The public code release lowers the barrier for groups wanting to reproduce or extend the method.
Future Directions
- Quantitative benchmarking: The abstract reports improvements over baselines but does not enumerate them. Independent replication across larger, multi-institutional cohorts would clarify how well the approach generalizes.
- Dose-level safety margins: A mean target dose error of 1.69% is reported as clinically acceptable, but organ-at-risk dose accuracy and worst-case errors are not addressed in the abstract. These are the figures that typically govern clinical adoption.
- Anatomical variability: The abstract names anatomical variability as a core challenge. Extending evaluation to unusual anatomy, implants, or post-surgical cases would test the method's robustness.
- Prospective clinical validation: Moving from retrospective comparison against existing CTs to prospective MRI-only planning studies would establish whether the synthesis is reliable enough to replace CT in practice.
- Efficiency and integration: Making 3D transformer inference fast enough for routine clinical use, and integrating the model into existing treatment planning systems, remain open engineering questions.
Target Audience
This paper is most relevant to researchers and graduate students working on medical image synthesis, cross-modality translation, or transformer architectures for 3D volumetric data. It is also directly useful to medical physicists and radiation oncologists evaluating MRI-only planning workflows, and to clinical software engineers assessing whether synthetic CT generation is mature enough for deployment. Readers without background in either deep learning or radiotherapy would find the terminology dense, though the clinical motivation is accessible.
Authors’ abstract
MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use for dose calculation. As a result, current radiotherapy workflows rely on combined MRI and CT acquisitions, increasing registration uncertainty and procedural complexity. Synthetic CT generation enables MRI only planning but remains challenging due to nonlinear MRI-CT relationships and anatomical variability. We propose Parallel Swin Transformer-Enhanced Med2Transformer, a 3D architecture that integrates convolutional encoding with dual Swin Transformer branches to model both local anatomical detail and long-range contextual dependencies. Multi-scale shifted window attention with hierarchical feature aggregation improves anatomical fidelity. Experiments on public and clinical datasets demonstrate higher image similarity and improved geometric accuracy compared with baseline methods. Dosimetric evaluation shows clinically acceptable performance, with a mean target dose error of 1.69%. Code is available at: https://github.com/mobaidoctor/med2transformer.