Skip to content
AI.info

Research

Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research

Overview Research area: Federated learning for medical image-to-image translation (MRI-to-CT synthesis), framed through the lens of Green AI and health equity. Technical level: Intermediate. The reade

Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research
arXiv
2512.03054
Published
2025-11-25
Authors
Ciro Benito Raggio, Lucia Migliorelli, Nils Skupien, Mathias Krohmer Zabaleta, Oliver Blanck, Francesco Cicone, Giuseppe Lucio Cascini, Paolo Zaffino, Maria Francesca Spadea

AI summary

Overview

Research area: Federated learning for medical image-to-image translation (MRI-to-CT synthesis), framed through the lens of Green AI and health equity.

Technical level: Intermediate. The reader should be comfortable with deep learning concepts such as encoder-decoder networks, weight aggregation, and transfer/convergence behavior, though the paper's core idea is described in accessible terms.

Scope (one sentence): This paper proposes a server-side, patience-based adaptive encoder-freezing mechanism that runs inside a federated MRI-to-CT training pipeline and measures whether it cuts training time, energy use, and CO2eq emissions without degrading synthetic CT quality across five encoder-decoder architectures.

What This Paper Is About

Federated learning lets hospitals train a shared deep learning model without moving patient data, which should broaden access to AI in healthcare. In practice, however, federated training is computationally heavy, so centres with limited infrastructure may be unable to join, potentially concentrating AI benefits in well-resourced institutions. The authors' goal is to reduce the cost of the local training step by automatically freezing the encoder portion of encoder-decoder networks once the federated clients have effectively agreed on the encoder's weights, while checking that synthetic CT quality is preserved.

Key Contributions

  1. A patience-based adaptive encoder-freezing strategy for federated learning. The central server tracks the relative percentage difference (ρ%) of the aggregated encoder weights between consecutive rounds, computed as the Mean Absolute Error between corresponding encoder layers. Freezing is triggered only when ρ% stays below a user-set threshold τ for N consecutive rounds (the patience parameter), preventing freezing on transient fluctuations.
  2. A Green AI-guided evaluation framework. The authors instrument the whole federation with the CodeCarbon library to report training time, total energy consumption (kWh), and CO2eq emissions (kg) alongside conventional image-quality metrics, treating efficiency and environmental cost as first-class outcomes rather than afterthoughts.
  3. A cross-architecture validation on MRI-to-CT conversion. Five encoder-decoder architectures (Simple U-Net, Li et al. 2024, Spadea/Pileggi et al., Fu et al., and Li et al. 2020) are tested in a four-client, one-server federated setup, with the frozen variant compared against an identically configured non-frozen counterpart over five repetitions each.
  4. An equity framing of federated optimization. The work positions reduced client-side training load as a mechanism for lowering the barrier to participation for resource-constrained clinical centres, tying a technical optimization to questions of fairness and justice in AI-driven healthcare.

Main Findings

  • Energy and emissions fell by up to roughly a quarter. Across architectures, total energy consumption and CO2eq emissions dropped between 9.1% and 23.2%; the average emissions reduction across all models was approximately 19.9%. Training-time improvements ranged from 9.4% (Spadea, Pileggi et al.) to 22.0% (Li et al.).
  • The Li et al. architecture gained the most. That model saw a 22.0% reduction in training time (10.11 ± 0.38 h to 7.89 ± 0.22 h), a 23.2% emissions reduction (1.51 ± 0.04 kg to 1.16 ± 0.03 kg CO2eq), and a 23.0% energy reduction (3.96 ± 0.10 kWh to 3.05 ± 0.07 kWh).
  • Simple U-Net, Fu et al., and Li et al. 2020 clustered around 16–17%. Simple U-Net: 16.8% faster training (13.40 ± 0.76 h to 11.15 ± 0.64 h), 16.7% lower emissions, 16.4% lower energy. Fu et al.: 16.2% faster, 17.1% lower emissions, 16.7% lower energy. Li et al. 2020: 16.6% faster, 16.8% lower emissions, 16.7% lower energy.
  • The deepest model improved least. Spadea, Pileggi et al. had the longest absolute training time and the smallest relative gain, at 9.4% (19.22 ± 0.75 h to 17.42 ± 0.68 h), with emissions and energy reductions of approximately 9%.
  • Trainable parameter counts dropped sharply. For example, Simple U-Net went from 1.727 × 10^7 to 7.857 × 10^6 parameters, and Li et al. 2020 from 2.682 × 10^7 to 9.024 × 10^6.
  • Per-epoch client-side savings were substantial. On Centre A, Simple U-Net fell from 21.07 ± 0.65 min to 14.38 ± 0.43 min per local epoch (31.8% reduction) and Li et al. from 15.59 ± 0.53 min to 10.05 ± 0.31 min (35.5%). The largest absolute reductions were on Centre D for Spadea, Pileggi et al. (44.03 ± 0.79 to 35.77 ± 0.59 min) and Fu et al. (38.50 ± 1.90 to 27.71 ± 1.35 min).
  • Conversion quality was preserved. On 23 unseen patients from the held-out Centre E, three of five architectures (Simple U-Net, Fu et al., Li et al. 2020) showed no statistically significant MAE difference (p > 0.05, t-test). Li et al. and Spadea, Pileggi et al. showed statistically significant improvements (p < 0.001), but the absolute MAE reduction of approximately 2–3 HU is far below clinically meaningful thresholds, where deviations greater than 20–50 HU are typically needed to affect tissue discrimination or dose distributions.
  • PSNR and SSIM were essentially unchanged. Median values sat near 26.3–26.6 for PSNR and 0.89 for SSIM in both settings, with only minor stochastic fluctuations.
  • A 10% threshold was too aggressive. In the observational study with patience N = 3, τ = 10% was reached after roughly 5–6 rounds but degraded performance across models, whereas τ = 5% was reached after 7–10 rounds and preserved performance, so τ = 5%, N = 3 was used in the final experiments.

Methodology in Plain English

Four clinical centres (A–D) each hold a private brain MRI/CT dataset and act as federated clients; a fifth institution (E) acts as the server and provides a completely unseen test set. Centres A–C contributed private data from the US, Italy, and Germany; centres D and E were single institutions drawn from the SynthRAD Grand Challenge 2023. Datasets contained 15, 14, 21, 29, and 23 cases respectively, with differing voxel spacings and image sizes (see Table 1 of the paper). Each centre preprocesses its own data locally (N4 bias field correction, rescaling and normalisation to [0,1], resizing/cropping/padding to 256 × 256 × 256, and spatial augmentation) so no raw images ever leave the site.

Training follows the earlier FedSynthCT-Brain protocol: PyTorch and MONAI for the models, the Flower framework for federation, batch size 8, learning rate 10^-4, one local epoch per client per round, the Adam optimizer, a Random Multi-2D training scheme with a voting strategy across anatomical planes at test time, and FedAvg combined with FedProx (proximal term of 3). Every experiment runs for a fixed 25 communication rounds and is repeated five times. Everything runs in a container with 32 GB RAM, one NVIDIA A100 80GB GPU, and 16 CPU cores, which supports the whole simulated federation on a single physical node; the authors note that an individual client in a real distributed deployment would not need comparable resources.

The novelty sits on the server. After each round the server measures how much the aggregated encoder weights moved relative to the previous round, expressed as ρ%. When ρ% stays under τ = 5% for N = 3 consecutive rounds, the server signals clients to freeze the encoder. Because frozen encoder layers no longer require back-propagation or aggregation, local epochs get faster. The decoder is deliberately never frozen, so clients retain the capacity to adapt to their own data distributions. CodeCarbon tracks time, energy, and CO2eq for every run, and image similarity is measured with MAE, PSNR, and SSIM against ground-truth CT. The comparison baseline is exactly the same pipeline with freezing disabled.

Why This Matters

Impact on research. The paper reframes federated learning optimization as a sustainability and equity problem rather than purely a communication-bandwidth problem. Prior freezing work (for example SmartFreeze, FedPT, and the method of Malan et al.) targeted communication cost or memory, and the authors argue those approaches were not aligned with Green AI principles or validated on complex clinical encoder-decoder architectures. This work extends layer freezing to a deep medical image translation task, adds an explicit patience criterion to distinguish transient noise from real convergence, and reports environmental cost alongside accuracy — a reporting pattern the authors position as groundwork for future federated evaluation frameworks.

Real-world applications:

  • MRI-only radiotherapy planning. Synthetic CT generated from MRI can replace the planning CT scan, reducing additional radiation exposure and streamlining workflows, provided the synthetic images remain accurate.
  • Federated model development across unequal hospital networks. Smaller or less well-equipped centres could join a federation using hardware that would otherwise be insufficient, without sharing patient data.
  • Carbon accounting for clinical AI. The energy and CO2eq tracking approach offers a template for hospitals and vendors that must report or reduce the environmental footprint of the models they train and deploy.
  • Privacy-preserving multi-institution collaboration. The setup allows institutions in different countries and regulatory regimes to contribute data indirectly to a single model while keeping raw images on site.

Industry relevance. Federated learning platform vendors, medical device and radiotherapy software companies, hospital consortia, and cloud providers all have a stake in making federated training cheap enough for the least-resourced participant, since a federation is only as useful as the number of sites that can realistically join it. The combination of a lightweight server-side rule and measurable energy savings is directly deployable, and the Green AI framing responds to growing institutional pressure to document the environmental cost of machine learning.

Future Directions

  • Extending freezing to other parts of the network, or testing where it should stop. The authors deliberately restricted freezing to the encoder, arguing that freezing the decoder would harm the global model's ability to accommodate local data heterogeneity. Whether a partial or scheduled decoder freezing could work is not explored here.
  • Validating the method beyond MRI-to-CT and beyond this four-client configuration. The paper does not report results for other image translation tasks, other modalities, other federation sizes, or non-brain anatomies.
  • Moving from a single-node simulation to a genuinely distributed deployment. The entire federation was run on one A100 host in this study; real-world behaviour over heterogeneous networks and hardware is not reported.
  • Per-client statistical validation. The authors state that statistical testing across individual clients was precluded by the limited number of test cases available per site, so client-level performance effects remain an open question. The discussion section of the provided content also ends mid-sentence and does not present a formal limitations or future work section, so further stated directions are not reported.

Target Audience

This paper is most valuable to federated learning researchers working on efficiency and client-incentive problems; medical imaging and radiotherapy researchers interested in MRI-to-CT synthesis; and Green AI or sustainable computing researchers looking for concrete, clinically motivated case studies. It also speaks to health equity and policy audiences concerned with whether AI infrastructure requirements widen or narrow disparities between institutions, and to engineers building federated platforms who need practical, low-overhead strategies that reduce the compute burden on participating sites.

Authors’ abstract

Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep learning (DL) models, even with limited data. However, the significant resource requirements of FL often exclude centres with limited computational infrastructure, further widening existing healthcare disparities. To address this issue, we propose a Green AI-oriented adaptive layer-freezing strategy designed to reduce energy consumption and computational load while maintaining model performance. We tested our approach using different federated architectures for Magnetic Resonance Imaging (MRI)-to-Computed Tomography (CT) conversion. The proposed adaptive strategy optimises the federated training by selectively freezing the encoder weights based on the monitored relative difference of the encoder weights from round to round. A patience-based mechanism ensures that freezing only occurs when updates remain consistently minimal. The energy consumption and CO2eq emissions of the federation were tracked using the CodeCarbon library. Compared to equivalent non-frozen counterparts, our approach reduced training time, total energy consumption and CO2eq emissions by up to 23%. At the same time, the MRI-to-CT conversion performance was maintained, with only small variations in the Mean Absolute Error (MAE). Notably, for three out of the five evaluated architectures, no statistically significant differences were observed, while two architectures exhibited statistically significant improvements. Our work aligns with a research paradigm that promotes DL-based frameworks meeting clinical requirements while ensuring climatic, social, and economic sustainability. It lays the groundwork for novel FL evaluation frameworks, advancing privacy, equity and, more broadly, justice in AI-driven healthcare.

Read the original paper