Skip to content
AI.info

Research

Med-MMFL: A Multimodal Federated Learning Benchmark in Healthcare

Overview Research area: Federated learning (FL) benchmarks for multimodal machine learning in healthcare, spanning medical image analysis (MRI, X-ray, pathology), ECG, electronic health records, radio

Med-MMFL: A Multimodal Federated Learning Benchmark in Healthcare
arXiv
2602.04416
Published
2026-02-04
Authors
Aavash Chhetri, Bibek Niroula, Pratik Shrestha, Yash Raj Shrestha, Lesley A Anderson, Prashnna K Gyawali, Loris Bazzani, Binod Bhattarai

AI summary

Overview

  • Research area: Federated learning (FL) benchmarks for multimodal machine learning in healthcare, spanning medical image analysis (MRI, X-ray, pathology), ECG, electronic health records, radiology text, and vision-language tasks.
  • Technical level: Intermediate. The paper assumes familiarity with federated learning concepts such as client/server aggregation, non-IID data partitioning, and Dirichlet-based label skew, as well as standard multimodal model architectures.
  • Scope: The paper introduces Med-MMFL, described as the first comprehensive benchmark for medical multimodal federated learning, covering five datasets with 2 to 4 modalities each, 10 unique medical modalities, six FL algorithms, four task types, and three partitioning strategies, with a public code release at https://github.com/bhattarailab/Med-MMFL-Benchmark.

What This Paper Is About

Training multimodal clinical models is difficult because medical data is fragmented across institutions and cannot be freely shared due to privacy regulations. Federated learning solves part of this by training models locally and aggregating them on a server, but existing medical FL benchmarks are mostly unimodal or bimodal and cover few tasks, making it hard to compare methods fairly. Med-MMFL addresses this by assembling a standardized, reproducible benchmark that evaluates six FL algorithms across diverse medical modalities, tasks, and realistic federation scenarios.

Key Contributions

  1. A unified MMFL benchmark. Med-MMFL integrates 6 distinct FL algorithms, 4 task types (segmentation, classification, modality alignment/retrieval, and VQA), 3 data partitioning strategies, and 5 medical datasets with varying degrees of multimodality ranging from 2 to 4 modalities.
  2. Generalization of existing algorithms to more than two modalities. The authors adapt MOON into m-MOON, which performs modality-wise contrastive alignment between local and global representations for each modality, and generalize CreamFL into CreamMFL, extending its contrastive regularizers to pairwise inter-modal and intra-modal losses plus global-local contrastive aggregation using the available modalities.
  3. Coverage beyond prior benchmarks. Compared to NIID-Bench, FLamby, FedLLM-Bench, FedMultimodal, and FedVLMBench, Med-MMFL reports the largest counts in the paper's comparison table for modality range (2 to 4), unique medical modalities (10), multimodal medical datasets (5), distinct FL algorithms (6), and evaluation task types (4), while supporting real-world, synthetic IID, and synthetic non-IID partitioning.
  4. Public release for reproducibility. The complete benchmark implementation, including all dataset processing and partitioning pipelines, is released openly.

Main Findings

  • No single algorithm wins everywhere. Across the benchmark, results show that no one FL algorithm consistently outperforms the others across all experimental configurations.
  • FedProx, FedAvg, and SCAFFOLD are the most consistent top performers, though in different contexts. The paper reports that FedProx achieves the highest performance in 7 out of 16 settings, particularly excelling in non-IID scenarios, while FedAvg and SCAFFOLD are strongest primarily in IID settings.
  • FedNova dominates brain tumor segmentation. Although FedNova does not rank first or second in most benchmarks, on Fed-BraTS-GLI2024 it achieves the best results in 4 settings and the second-best in 2 settings, totaling top performances in 6 out of 7 configurations.
  • Per-dataset leaders differ. On Fed-BraTS-GLI2024, FedNova performs best in 4 out of 7 settings. On Fed-MIMIC-CXR-JPG, FedAvg and FedProx lead in 2 settings each, together covering 4 of 6 configurations. On Fed-Symile-MIMIC, FedAvg dominates 3 settings and FedProx 2, while SCAFFOLD and FedNova perform notably worse. On Fed-PathVQA, SCAFFOLD achieves the best results in 2 of 3 settings, and on Fed-EHRXQA, FedProx leads in 2 out of 3 cases.
  • Overall leaders across datasets. FedAvg and FedProx rank highest in 3 out of 5 datasets in this experimental setup.
  • Partitioning strategy matters unevenly. Fed-BraTS-GLI2024 and Fed-MIMIC-CXR-JPG remain largely consistent across splits, except CreamMFL, which varies by up to 5–6% in the 3-client setting. Fed-Symile-MIMIC shows a significant drop from IID to non-IID configurations across all algorithms, attributed to contrastive learning's dependence on diverse batch samples. Fed-PathVQA and Fed-EHRXQA remain robust to data heterogeneity.
  • Centralized training is only modestly better in most cases. Differences between centralized and federated training are typically less than 1–1.5%. The exception is MIMIC-CXR-JPG, where centralized training achieves nearly 7% higher accuracy than the best federated algorithm. The authors also report cases where federated algorithms outperform the centralized baseline, which they attribute to implicit regularization from heterogeneous client updates.
  • The hardest dataset is multimodal retrieval on Symile-MIMIC. Centralized accuracy reaches 41.370, while the federated evaluation range is (9.2672, 38.147), with the lowest value coming from FedNova.
  • Centralized reference points (all clients' data pooled): BraTS-GLI2024 DSC 84.900 (150 epochs), MIMIC-CXR-JPG AUC 92.380 (40 epochs), Symile-MIMIC Acc 41.370 (50 epochs), PathVQA F1 87.124 (60 epochs), EHRXQA F1 52.385 (30 epochs).
  • SCAFFOLD instability on Symile-MIMIC. The authors note that gradient scaling was applied to control gradient correction in SCAFFOLD for that dataset, and hypothesize that degradation stems from instability introduced by frequent gradient corrections or cumulative gradient-based normalization.

Methodology in Plain English

The authors selected five publicly available medical datasets and converted each into a federated version. Where the data contains information about which institution contributed each sample, they create a natural partition by assigning each contributing center to a client. Where that information is absent, they create synthetic partitions: an IID split where labels are spread roughly evenly across clients (achieved by sampling from a Dirichlet distribution with a concentration parameter approaching infinity), and non-IID splits using Dirichlet sampling with α = 0.8 and α = 0.2 to control how skewed each client's label distribution becomes. For datasets without class labels suitable for partitioning, they derive pseudo-classes by clustering (for example, clustering class-proportion vectors from segmentation masks, or clustering question embeddings produced by BioMedClip).

Each dataset is paired with a task and a metric: brain tumor sub-region segmentation with Dice Score Coefficient using an RFNet baseline; 14-label chest X-ray classification with macro AUC using a multimodal classifier; chest X-ray / ECG / blood-lab alignment with zero-shot retrieval accuracy using ResNet-50 for CXR, ResNet-18 for ECG, and a three-layer MLP for labs; closed-ended yes/no pathology VQA with F1 using a BLIP baseline; and single-patient EHR + chest X-ray VQA with Token Overlap F1 using BLIP's generative decoder.

Six FL algorithms are evaluated: FedAvg, FedProx, SCAFFOLD, m-MOON (the authors' multimodal extension of MOON), FedNova, and CreamMFL (the authors' extension of CreamFL). For CreamMFL, a separate model is trained on public data as a virtual client so that client configurations remain consistent for fair comparison. Experiments use cross-silo settings with full client participation, the Adam optimizer for local optimization except FedNova (which instead uses SGD with momentum and an independently tuned learning rate), identical training schedules across algorithms, three independent runs with different random seeds, and NVIDIA A100-PCIE-40GB GPUs. The number of communication rounds and local epochs, and the full hyperparameter set, are reported as being in the supplementary material rather than in the visible portion of the paper.

Why This Matters

Impact on research. Medical FL research has been constrained by fragmented, mostly unimodal benchmarks that make cross-paper comparison unreliable. Med-MMFL widens evaluation to 2–4 modalities per dataset, four task types, and three partitioning strategies, giving the community a shared protocol and open implementation. Extending MOON and CreamFL to more than two modalities provides a concrete template for how old algorithms can be fairly compared in multimodal settings.

Real-world applications.

  • Multi-hospital brain tumor segmentation from four MRI sequences, where each hospital holds its own scanner data and cannot share it.
  • Chest X-ray interpretation alongside radiology reports across institutions, supporting 14-label multi-label diagnostic classification.
  • Cross-modal retrieval linking chest X-rays to ECG recordings and blood laboratory measurements, useful for aligning records that were never captured together.
  • Clinical question answering over patient-specific electronic health records plus imaging, including pathology VQA that mirrors American Board of Pathology-style reasoning.

Industry relevance. Hospital networks, medical imaging vendors, and health-AI companies that need to train across institutional silos will benefit from knowing which aggregation strategy holds up under real-world data skew. The finding that FedProx is comparatively robust under non-IID conditions, while FedNova is strong for segmentation, gives practical guidance for deployment choices. The observation that federated training sometimes beats centralized training also matters commercially, since it suggests privacy-preserving training need not always cost accuracy.

Future Directions

  • Personalized federated learning. The authors explicitly state that personalized FL lies outside the scope of this work and is open for future exploration.
  • Broader algorithm coverage. Several state-of-the-art methods (the paper mentions FedDyn, FedOpt-family variants such as FedAvgM, FedAdam, FedAdagrad, and FedYogi) are not directly evaluated, and the authors note the algorithm set in existing benchmarks is narrow.
  • Fusion strategy design. Med-MMFL focuses on aggregating client updates rather than designing or evaluating multimodal fusion strategies; a systematic study of fusion methods is stated as beyond the scope of this work.
  • Extending beyond the current dataset and modality set. The benchmark covers 5 datasets and up to 4 modalities, so the paper leaves open whether its findings transfer to other clinical data types and to cross-device rather than cross-silo federation settings.

Target Audience

This paper is most valuable to federated learning researchers, medical AI practitioners, and machine learning engineers building privacy-preserving multi-institutional systems. It also serves benchmark designers and reproducibility-focused researchers who need standardized evaluation protocols, and graduate students entering either FL or multimodal medical ML who want a single reference point for datasets, baselines, and partitioning conventions. Clinicians and hospital IT decision-makers evaluating whether federated training is ready for deployment

Authors’ abstract

Federated learning (FL) enables collaborative model training across decentralized medical institutions while preserving data privacy. However, medical FL benchmarks remain scarce, with existing efforts focusing mainly on unimodal or bimodal modalities and a limited range of medical tasks. This gap underscores the need for standardized evaluation to advance systematic understanding in medical MultiModal FL (MMFL). To this end, we introduce Med-MMFL, the first comprehensive MMFL benchmark for the medical domain, encompassing diverse modalities, tasks, and federation scenarios. Our benchmark evaluates six representative state-of-the-art FL algorithms, covering different aggregation strategies, loss formulations, and regularization techniques. It spans datasets with 2 to 4 modalities, comprising a total of 10 unique medical modalities, including text, pathology images, ECG, X-ray, radiology reports, and multiple MRI sequences. Experiments are conducted across naturally federated, synthetic IID, and synthetic non-IID settings to simulate real-world heterogeneity. We assess segmentation, classification, modality alignment (retrieval), and VQA tasks. To support reproducibility and fair comparison of future multimodal federated learning (MMFL) methods under realistic medical settings, we release the complete benchmark implementation, including data processing and partitioning pipelines, at https://github.com/bhattarailab/Med-MMFL-Benchmark .

Read the original paper