Skip to content
AI.info

Research

FedOnco-Bench: A Reproducible Benchmark for Privacy-Aware Federated Tumor Segmentation with Synthetic CT Data

Overview Research area: Federated learning for medical image segmentation, with a focus on privacy leakage (membership inference) and differential privacy; uses synthetic oncologic CT data. Technical

FedOnco-Bench: A Reproducible Benchmark for Privacy-Aware Federated Tumor Segmentation with Synthetic CT Data
arXiv
2511.00795
Published
2025-11-02
Authors
Viswa Chaitanya Marella, Suhasnadh Reddy Veluru, Sai Teja Erukude

AI summary

Overview

Research area: Federated learning for medical image segmentation, with a focus on privacy leakage (membership inference) and differential privacy; uses synthetic oncologic CT data.

Technical level: Intermediate. The core ideas (federated averaging, non-IID client data, DP-SGD noise injection, membership inference AUC) are explained in the paper, but readers benefit from prior familiarity with federated learning and segmentation metrics.

Scope: The paper introduces FedOnco-Bench, a reproducible open-source benchmark that simulates five federated clients on a synthetic CT tumor-segmentation dataset and measures both segmentation accuracy (Dice, cross-entropy) and privacy risk (membership-inference attack AUC) for FedAvg, FedProx, FedBN, and FedAvg with DP-SGD.

What This Paper Is About

Federated learning lets hospitals train a shared model without exchanging patient images, but trained models can still leak whether a specific patient's scan was used in training, and client data is often heterogeneous. FedOnco-Bench provides a controlled, fully synthetic testbed where segmentation accuracy and privacy leakage can be measured side by side, so different federated methods can be compared on the same footing. The goal is to quantify the privacy-utility tradeoff that medical federated learning faces rather than reporting accuracy alone.

Key Contributions

  1. Synthetic federated dataset: A synthetic CT dataset of 5,000 annotated 2D axial slices (256 × 256) with one or more tumor regions each, generated with a diffusion-based model akin to DiffGuard, deliberately split non-IID across five simulated clients (for example, Client 1 predominantly large tumors, Client 2 smaller lesions, Client 3 noisy images).
  2. Privacy-preserving FL baselines: Implementation and evaluation of FedAvg, FedProx, and FedBN alongside a DP-enhanced FedAvg (FedAvg + DP-SGD), with a secure aggregation protocol assumed so the server sees only the sum of updates.
  3. Dual metrics and evaluation: Segmentation performance measured by Dice coefficient and pixel-wise cross-entropy loss, and privacy leakage measured by the AUC of a black-box membership inference attack, with per-round training curves for both.
  4. Reproducibility: Public release of code and data generation scripts at https://github.com/viswachaitanyamarella/FedOnco-Bench.

Main Findings

  • Federated accuracy approaches centralized accuracy: FedAvg and FedBN both reach a mean Dice of approximately 0.85, and FedProx reaches 0.84, while the centralized model trained on pooled data for 500 epochs reaches 0.88. The differences among FedAvg, FedBN, and FedProx (±0.01 Dice) are described as not statistically significant.
  • Cross-entropy loss follows the same ordering: FedAvg 0.34, FedBN 0.35, FedProx 0.36, centralized 0.30.
  • Non-private federated models leak membership: Membership inference AUC is 0.72 for FedAvg, 0.70 for FedBN, and 0.68 for FedProx, all above the 0.5 random-guess level.
  • Centralized training leaks about as much as FedAvg: The centralized model has an MI AUC of 0.72, the same as FedAvg, suggesting overfitting is a concern even without federation.
  • DP-SGD sharply reduces leakage at an accuracy cost: FedAvg + DP-SGD reaches an MI AUC of 0.25, near random guessing, while Dice drops to 0.79 and cross-entropy loss rises to 0.42 (a roughly 6-point Dice drop from 0.85 and a CE increase from 0.34 to 0.42).
  • Convergence speed differs by method: FedAvg and FedBN plateau near 0.85 by round 60, FedProx improves gradually to 0.84 by round 100, and DP-SGD shows slower, noisier improvement, peaking at 0.79.
  • Leakage appears early: MI risk for FedAvg and FedBN rises and stabilizes around 0.72 during training, FedProx saturates lower near 0.68, and DP-SGD stays flat at 0.25 throughout, suggesting most leakage occurs early when the model memorizes data.
  • The local (non-federated) baseline is weak and leaky: Independent per-client models achieve a mean Dice of approximately 0.70 with MI risk of approximately 0.80, and are therefore excluded from the main results table.

Methodology in Plain English

The authors simulate a cross-silo federated system with one central server and five hospitals (clients). Rather than using real patient scans, they generate synthetic CT slices with a diffusion-based generative model, annotate tumor regions, and intentionally give each client a different tumor profile so the data is non-IID (for example, mostly large tumors at one client, small nodules at another, noisy images at a third). Clients 1–3 receive 1,000 unique training slices each and clients 4–5 receive 500 each, producing unbalanced data; each client keeps 80% of its local data for training and 20% as a local validation set that is never shared, and a separate 1,000-image global test set is used for final evaluation.

The segmentation model is a small 2D U-Net (two down-blocks, two up-blocks, skip connections, batch normalization and ReLU, roughly 1.2M parameters) trained with pixel-wise cross-entropy loss and scored with Dice. In each of 100 federated rounds, the server broadcasts the global model, each client trains locally for one epoch, and the server averages the updates (fed by a secure aggregation assumption so individual updates stay hidden). FedProx adds a proximal penalty with μ = 0.01 to keep local models close to the global one; FedBN keeps batch normalization parameters local and aggregates only convolutional weights. The DP variant clips each gradient to an L2 norm of C = 1.0 and adds Gaussian noise with σ = 1.2, which the paper describes as approximating a moderate privacy budget (ε < 10 per round, though ε is not computed explicitly). Training uses SGD with batch size 16, learning rate 0.01 decayed by 0.1 at round 70, momentum 0.9, weight decay 1 × 10⁻⁴, and each method is run three times with different seeds, reporting mean ± standard deviation.

Privacy is measured with a black-box membership inference attack: the attacker gets predictions on 500 training samples (members) and 500 unseen samples (non-members), trains a shadow model of the same architecture on a separate shadow dataset of 1,000 synthetic images distributed across five shadow clients, trains an attack classifier on the shadow model's outputs, and reports the AUC of membership classification, where 0.5 is random guessing and 1.0 is full leakage.

Why This Matters

The paper moves federated medical imaging evaluation beyond accuracy alone, providing a common, fully synthetic, reproducible framework in which privacy leakage is an explicit reported metric. Because the data are synthetic, the benchmark can be shared publicly without patient privacy concerns, which is not possible with real clinical datasets.

Real-world applications:

  • Hospital consortia training tumor segmentation models across institutions without pooling patient scans, where the benchmark shows what accuracy and privacy costs to expect from each aggregation strategy.
  • Selecting a privacy operating point for deployment: teams that cannot tolerate membership inference can choose DP-SGD and accept lower Dice, while teams prioritizing accuracy can see the leakage risk of plain FedAvg.
  • Regulatory and compliance discussions about whether a deployed federated model can reveal whether a patient participated in training, using MI AUC as a concrete, reportable measure.
  • Method development and ablation studies, since the benchmark supplies fixed baselines for new federated optimization or privacy mechanisms.

Industry relevance: the benchmark is relevant to medical imaging vendors, radiology AI developers, and cloud or platform providers offering federated training infrastructure for healthcare, because it packages the privacy-utility tradeoff into a comparable, open-source testbed rather than a single accuracy number.

Future Directions

  • Extend beyond 2D synthetic CT: the authors suggest adding other modalities such as synthetic MRI for brain tumor segmentation or digital pathology imaging for cell structure classification, and note that real 3D CTs add complexity such as texture and artefacts that the 2D slices do not capture.
  • Explore more privacy mechanisms and budgets: the paper tested only one DP noise scale and did not trace the full privacy curve by varying ε; it suggests varying the privacy budget, using user-level differential privacy accounting across training rounds, and investigating homomorphic encryption and split learning.
  • Test stronger and broader attacks: only black-box membership inference was studied; white-box attacks, model inversion, and attribute inference were not considered.
  • Try more federated algorithms and realistic deployment conditions: the authors propose momentum-based FedAvgM and personalized methods such as FedPer, and note that their setup assumes all clients participate in every round, unlike real systems with partial participation and communication constraints.

Target Audience

Researchers and practitioners in federated learning, medical image analysis, and privacy-preserving machine learning; graduate students looking for a reproducible starting point for federated segmentation experiments; and clinical AI or compliance teams evaluating the privacy-utility tradeoff of federated tumor segmentation before deployment.

Authors’ abstract

Federated Learning (FL) allows multiple institutions to cooperatively train machine learning models while retaining sensitive data at the source, which has great utility in privacy-sensitive environments. However, FL systems remain vulnerable to membership-inference attacks and data heterogeneity. This paper presents FedOnco-Bench, a reproducible benchmark for privacy-aware FL using synthetic oncologic CT scans with tumor annotations. It evaluates segmentation performance and privacy leakage across FL methods: FedAvg, FedProx, FedBN, and FedAvg with DP-SGD. Results show a distinct trade-off between privacy and utility: FedAvg is high performance (Dice around 0.85) with more privacy leakage (attack AUC about 0.72), while DP-SGD provides a higher level of privacy (AUC around 0.25) at the cost of accuracy (Dice about 0.79). FedProx and FedBN offer balanced performance under heterogeneous data, especially with non-identical distributed client data. FedOnco-Bench serves as a standardized, open-source platform for benchmarking and developing privacy-preserving FL methods for medical image segmentation.

Read the original paper