Research
Self-Supervised Slice-to-Volume Reconstruction with Gaussian Representations for Fetal MRI
Overview Research area: Medical image analysis and computer vision, specifically slice-to-volume reconstruction (SVR) for fetal brain MRI, combining 3D Gaussian representations with self-supervised le

- arXiv
- 2601.22990
- Published
- 2026-01-30
- Authors
- Yinsong Wang, Thomas Fletcher, Xinzhe Luo, Aine Travers Dineen, Rhodri Cusack, Chen Qin
AI summary
Overview
Research area: Medical image analysis and computer vision, specifically slice-to-volume reconstruction (SVR) for fetal brain MRI, combining 3D Gaussian representations with self-supervised learning.
Technical level: Advanced. The paper assumes familiarity with 3D Gaussian Splatting, implicit neural representations, rigid registration, and MRI acquisition modelling.
Scope: The paper proposes GaussianSVR, a self-supervised framework that represents a fetal brain volume as a set of 3D Gaussians and jointly optimizes those Gaussians and slice-wise rigid transformations across multiple resolution levels, using a simulated forward slice acquisition model instead of ground-truth volumes.
What This Paper Is About
Fetal MRI is acquired as stacks of fast 2D slices, but unpredictable fetal motion leaves the slices misaligned with respect to one another, so cross-sectional views do not directly show true 3D brain structure. Slice-to-volume reconstruction tries to recover a single high-resolution 3D volume from these motion-corrupted stacks. The paper's goal is to do this without ground-truth volumes for training, which learning-based SVR methods have previously required and which are inaccessible in practice.
Key Contributions
-
A 3D Gaussian representation for SVR. The authors state they are the first to propose an SVR framework based on 3D Gaussian representation, replacing the voxel grid used by conventional optimization methods and the globally parameterized implicit neural representation used by NeSVoR.
-
A self-supervised multi-resolution training strategy. Slice-wise transformations and Gaussian parameters are jointly optimized hierarchically across resolution levels, using a simulated forward slice acquisition model to compare reconstructed slices against acquired slices.
-
An adapted Gaussian parameterization for MRI. Following Li et al. [9], the appearance-related parameters of standard 3D Gaussian Splatting (opacity alpha_j and spherical harmonics SH_j) are removed, and an intensity coefficient I_j is introduced to represent the MRI intensity value at each Gaussian center.
-
Empirical validation on the FeTA dataset. Experiments compare GaussianSVR against NiftyMIC, SVoRT, and NeSVoR, plus ablation studies isolating the multi-resolution and transformation-optimization components.
Main Findings
-
Highest quantitative accuracy on FeTA. Averaged over 30 test subjects, GaussianSVR reaches PSNR 28.19 dB (3.02), SSIM 0.9281 (0.0552), and NRMSE 0.0468 (0.0219). The next-best method, NeSVoR, reaches PSNR 25.58 dB (1.81), SSIM 0.8940 (0.0407), NRMSE 0.0536 (0.0105). The paper reports a 2.9% improvement in PSNR over NeSVoR.
-
Other baselines trail further behind. NiftyMIC scores PSNR 21.17 dB (1.95), SSIM 0.7653 (0.0559), NRMSE 0.0989 (0.0234). SVoRT scores PSNR 23.98 dB (2.65), SSIM 0.8209 (0.0618), NRMSE 0.0905 (0.1227).
-
Statistical significance. Paired t-tests report p-value < 0.01 for GaussianSVR versus NiftyMIC, SVoRT, and NeSVoR on PSNR and SSIM. For NRMSE, NiftyMIC and SVoRT are marked significant, while NeSVoR is not marked.
-
Multi-resolution training matters. Removing it ("w/o low resolution") drops performance to PSNR 27.08 dB (3.89) and SSIM 0.9134 (0.0547). The authors attribute this to slice-wise transformations being more stable and converging more effectively at lower resolutions.
-
Joint transformation optimization matters more. Removing it ("w/o transformation optimization") causes a substantial drop to PSNR 22.86 dB (2.38) and SSIM 0.8148 (0.0752), which the authors interpret as evidence that joint optimization helps the model escape local minima and reach better global convergence.
-
Qualitative sharpness. In the qualitative comparison on a single subject, GaussianSVR is reported to reconstruct more fine-grained details and sharper detail than NeSVoR.
-
Spatial locality as the mechanism. The authors argue that 3D Gaussian kernels give spatially localized, independent primitives that adapt to complex anatomical structure while preserving global consistency, and that the smooth nature of the kernels provides implicit regularization of the reconstructed volume.
Methodology in Plain English
The target 3D brain volume is not stored as a grid of voxels. Instead it is stored as a collection of 3D Gaussian "blobs," each with a center position, a covariance (decomposed into a scaling matrix and a rotation matrix), and an intensity value. The intensity at any point in space is computed by summing the contributions of nearby Gaussians only, using a 99% confidence interval (mu_j +/- 3 sigma_j) to keep computation efficient.
To train without ground truth, the framework simulates how an MRI scanner would have produced each 2D slice. Applying an estimated transformation, an anisotropic Gaussian point-spread-function blur, and a down-sampling step to the current volume produces a "reconstructed stack." That stack is compared against the actually acquired stack using a loss combining an L1 data fidelity term, a differentiable structural similarity term, and a total variation regularizer on the volume.
Optimization runs coarse-to-fine. In the low-resolution stage the volume is downsampled by a factor of two, and both the Gaussian parameters and the slice transformations are optimized; this stage is intended to stabilize training and give a good initialization, because rigid motion is easier to estimate when fine detail is suppressed. The high-resolution stage then refines both parameter sets at full resolution to recover anatomical detail and improve alignment. The method is implemented in PyTorch on an NVIDIA A6000 Ada GPU with the Adam optimizer, and transformation parameters from a pretrained SVoRT model initialize the optimization. The mean learning rate decays from 2 x 10^-3 to 2 x 10^-6, with constant rates of 0.05, 0.005, and 0.001 for intensity, scaling, and rotation; transformation learning rates are 5 x 10^-4 for translation and 5 x 10^-5 for rotation.
Experiments use the FeTA dataset of T2-weighted fetal brain MR images. Thirty volumes were randomly selected as ground truths, registered to a fetal brain atlas, and resampled to 0.8 x 0.8 x 0.8 mm. Simulated 2D slices have 1 mm x 1 mm resolution, slice thickness between 2.5 and 3.5 mm, and size 128 x 128. For each subject, three stacks of 15-30 slices were simulated along orthogonal orientations, with fetal brain motion trajectories generated following prior work.
Why This Matters
Impact on research. The paper removes the dependency on ground-truth volumes or ground-truth transformations that limited prior learning-based SVR methods, and it introduces a representation (localized 3D Gaussians) that the authors claim beats both a conventional optimization method and an implicit neural representation on the same benchmark. It also extends Gaussian splatting, previously developed for natural-image novel-view synthesis, into volumetric MRI reconstruction and motion correction, an application the authors say had not been explored before.
Real-world applications:
- Recovering high-resolution 3D fetal brain volumes from routine motion-corrupted 2D MRI acquisitions, supporting studies of fetal brain development.
- Retrospective motion correction for any imaging protocol that acquires stacks of 2D slices with residual inter-slice motion.
- Adjacent medical use cases of Gaussian representations cited in the paper: CT reconstruction, surgical navigation, and surgical scene reconstruction.
- As a downstream building block for atlas registration and volumetric analysis pipelines in fetal imaging.
Industry relevance. Reconstruction quality improvements in PSNR, SSIM, and NRMSE on a public benchmark, plus the availability of code, make the approach directly testable in medical imaging software pipelines. The self-supervised design matters commercially because it avoids the annotation cost of paired ground-truth volumes. The paper does not report runtime or memory benchmarks against the baselines, so any efficiency claim beyond the stated motivation for multi-resolution training is not quantified.
Future Directions
-
Resolution-agnostic behaviour. The authors explicitly state that future work will explore the resolution-agnostic capability of GaussianSVR.
-
Single-stack reconstruction. They also name reconstruction from a single stack of slices as future work, which would remove the current reliance on three orthogonal stacks.
-
Reducing initialization dependence. GaussianSVR is initialized with transformation parameters from a pretrained SVoRT model, so removing that dependence on another trained model is an open direction implied by the setup.
-
Broadening evaluation. Current results cover 30 FeTA subjects with simulated motion and simulated slices; extension to more subjects, other fetal organs, and other acquisition settings is not reported.
Target Audience
Researchers and graduate students in medical image analysis, computer vision, and MRI physics who work on motion correction, super-resolution, or implicit and explicit neural scene representations. It is also relevant to engineers building fetal or general MRI reconstruction pipelines, and to readers already familiar with 3D Gaussian Splatting who want to see it applied to a volumetric medical reconstruction problem.
Authors’ abstract
Reconstructing 3D fetal MR volumes from motion-corrupted stacks of 2D slices is a crucial and challenging task. Conventional slice-to-volume reconstruction (SVR) methods are time-consuming and require multiple orthogonal stacks for reconstruction. While learning-based SVR approaches have significantly reduced the time required at the inference stage, they heavily rely on ground truth information for training, which is inaccessible in practice. To address these challenges, we propose GaussianSVR, a self-supervised framework for slice-to-volume reconstruction. GaussianSVR represents the target volume using 3D Gaussian representations to achieve high-fidelity reconstruction. It leverages a simulated forward slice acquisition model to enable self-supervised training, alleviating the need for ground-truth volumes. Furthermore, to enhance both accuracy and efficiency, we introduce a multi-resolution training strategy that jointly optimizes Gaussian parameters and spatial transformations across different resolution levels. Experiments show that GaussianSVR outperforms the baseline methods on fetal MR volumetric reconstruction. Code is available at https://github.com/Yinsong0510/GaussianSVR-Self-Supervised-Slice-to-Volume-Reconstruction-with-Gaussian-Representations.