Skip to content
AI.info

Research

Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data

Overview Research area: Multimodal federated learning (MMFL), representation learning, robustness to missing data. Technical level: Advanced. The paper combines a contrastive representation-learning d

Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data
arXiv
2510.22880
Published
2025-10-27
Authors
Duong M. Nguyen, Trong Nghia Hoang, Thanh Trung Huynh, Quoc Viet Hung Nguyen, Phi Le Nguyen

AI summary

Overview

  • Research area: Multimodal federated learning (MMFL), representation learning, robustness to missing data.
  • Technical level: Advanced. The paper combines a contrastive representation-learning design, a non-parametric clustering step for server aggregation, and a Lipschitz-based generalization bound.
  • Scope: The paper proposes PEPSY, a federated framework in which each client learns a "data-missing profile" of embedding controls that reconfigures a globally aggregated model to that client's own pattern of missing modalities and missing features, and evaluates it on the PTBXL and Sleep-EDF benchmarks.

Paper details: arXiv:2510.22880v1 [cs.LG], 27 Oct 2025. Authors: Duong M. Nguyen (University of Illinois Urbana-Champaign), Trong Nghia Hoang (Washington State University), Thanh Trung Huynh (VinUniversity), Quoc Viet Hung Nguyen (Griffin University), Phi Le Nguyen (Hanoi University of Science and Technology). Code is provided at https://github.com/nmduonggg/PEPSY.

What This Paper Is About

Standard federated learning assumes every client trains on the same set of feature modalities. In real deployments this is false in two ways at once: different clients may only have access to different subsets of modalities, and within the modalities a client does have, individual input features may be partly missing. When local models are trained on these different, incomplete views, they learn representations that do not line up with each other, so simply averaging them can destroy useful information. The paper's goal is to let the server aggregate a shared model while giving each client a way to reconfigure that shared representation to match its own missing-data situation, without sharing any raw data.

Key Contributions

  1. A new MMFL framework (PEPSY) with client-side embedding controls. Each client encodes its feature modalities, its data specifics, and its data-missing patterns into a local "data-missing profile" Ψ consisting of τ embedding controls. These controls act as reconfiguration signals that adapt the globally aggregated representation to the client's local context. The profiles are sent to the server, where they are aligned and aggregated, allowing clients with similar missingness profiles to collaborate.

  2. A theoretical analysis with an explicit performance bound. Theorem 3.1 bounds the expected deviation between the model's output when modalities are missing and when all modalities are present, in terms of the training loss L_ds, the number of missing modalities |S|, and the Lipschitz constant μ of the client model. The bound shows the deviation is driven down as L_ds is minimized during training.

  3. A non-parametric aggregation scheme for missingness profiles. Because clients can select different numbers of embedding controls, the paper treats profile aggregation as a non-parametric clustering problem, adopting PFPT to let the number of clusters adapt dynamically to the missingness level of the system. Ordinary neural components are still aggregated with FedAvg.

  4. Extensive empirical evaluation against five baselines. Experiments on PTBXL (12 modalities) and Sleep-EDF (5 modalities) across IID and Non-IID settings, multiple missing statistics, and mismatched train/test missing statistics, reporting up to a 36.45% performance improvement under severe data incompleteness.

Main Findings

  • Consistent superiority under matched missing statistics (IID). On PTBXL, differences are small when the missing degree is low (e.g., p_m = 0.2), but PEPSY keeps a significant advantage as missingness grows (e.g., p_m = 0.8). On EDF, PEPSY surpasses the baselines by up to 11.67% across all missing scenarios.

  • Best accuracy in 40/40 reported cases. The paper states that while most methods drop substantially, PEPSY remains robust and attains the highest accuracy in 40 out of 40 cases, attributed to the data-missing profile providing an informative reconfiguration signal.

  • Strong Non-IID gains. On PTBXL in the Non-IID setting, PEPSY surpasses FedMAC and other approaches by nearly 15.83% in lightly missing scenarios (p_m = 0.2), and reaches 64.69% accuracy at p_m/p_s = 0.8/0.8. It also leads on EDF under Non-IID conditions.

  • Robustness when train and test missing statistics differ. In Table 2, when clients have no missing data (p_m/p_s = 0.0/0.0), PEPSY achieves the highest accuracy in most testing missing scenarios, surpassing other baselines by an average of 3.45%. Under the challenging inter-client scenario (p_m/p_s = 0.5/1.0), PEPSY outperforms competitors by up to 14%.

  • Aggregation ablation. Table 3 compares FedAvg, FedProx, their probabilistic-alignment variants (SynFedProx, and SynFedAvg, which the paper says is used in PEPSY), and PEPSY across p_m/p_s settings of 0.2, 0.6, and 1.0. PEPSY and the synchronized variants generally outperform plain FedAvg and FedProx; the truncated text leaves the full discussion of this comparison incomplete.

  • Theory matches observation. Theorem 3.1 shows the stability of PEPSY under varying missing patterns depends on three factors: alignment quality of data-specific features (L_ds), the number of missing modalities (|S|), and the smoothness of the learned model (μ). When S → ∅ the bound converges to zero; in the worst case |S| = |M| − 1, it reduces to a quantity depending only on μ.

  • Visual evidence. t-SNE visualizations of modality representations (Figure 4, EDF dataset, Non-IID) and visualizations of global control embeddings over 500 training iterations (Figure 5b) are reported; the paper states the reduced distance between consecutive iterations indicates convergence and the variation shows embeddings capture different aspects from each client.

Methodology in Plain English

Each client decomposes its multimodal input into three kinds of information:

  1. Modality-specific features — shared learnable embeddings, one per modality, so they do not depend on the particular sample.
  2. Data-specific features — per-sample representations. For modalities the client actually has, these are mapped and normalized encodings; for missing modalities, they are reconstructed by averaging over the available modalities. A contrastive loss (L_ds) pulls features from the same instance's available modalities closer than features from different instances, so the representation preserves instance identity rather than the missing pattern.
  3. Missing-pattern features — the client queries its own data-missing profile Ψ (a pool of τ embedding controls) using the concatenated modality- and data-specific features as a query and the embedding controls as keys, scored by cosine similarity. Only the κ most relevant controls are selected (κ ≪ |Ψ|), and a regularization term R encourages this sparse selection. The averaged selected embeddings become the missing-pattern representation.

The three pieces are concatenated into a final representation for each modality. A second contrastive loss (L_rc) encourages the projected representations of the same instance to agree, acting as a reconfiguration signal that pushes data-missing embeddings toward data-complete form. A similarity-based attention mechanism fuses the per-modality representations into a cross-modal representation, which is combined with the original representation through a learned weighting. The full local objective combines the task loss, the two contrastive losses, and the sparsity regularizer: L = L_task + λ(L_ds + L_rc) − ηR, with λ and η controlling the contributions.

On the server, ordinary neural parameters are averaged with FedAvg, but the data-missing profiles cannot be averaged directly because clients learn them in arbitrary orders — the same embedding slot may mean different things for different clients. The server therefore treats profile aggregation as non-parametric clustering (using PFPT), grouping similar controllers and updating the global profile, whose size and complexity reflect the missingness level of the whole system. This repeats for T communication rounds.

The experiments simulate missingness with a binary mask matrix over modalities and samples. p_s is the ratio of samples with missing modalities, p_m is the ratio of missing modalities within those samples, and the missing degree is p_m × p_s. Each dataset is split 80% for training and 20% for testing, with the training portion distributed across K clients in both IID and Non-IID settings.

Why This Matters

  • Research impact. The paper targets a general MMFL setting that prior work addresses only in isolation — either varying modality sets without missing input features, or a shared modality set with missing features. The idea of learning an explicit, communicable "missingness profile" per client is a distinct alternative to imputation-based approaches, which the authors argue require access to all missing-data patterns and therefore cannot work federatively. It also provides a theoretical bound linking a training loss to robustness under missing modalities.

  • Real-world applications (the paper names wearable health monitoring, distributed environmental sensing, and smart infrastructure as domains with decentralized collection and frequent sensor failures):

    • Wearable health monitoring, where one device collects audio and another collects physiological signals, and recordings drop out intermittently.
    • Distributed environmental sensing, where sensor nodes differ in which measurements they capture.
    • Smart infrastructure, where deployed sensors fail or report partially.
    • Clinical or scientific settings where no pretrained foundation model spans all the required feature modalities, a limitation the paper explicitly raises for healthcare.
  • Industry relevance. The framework requires no raw data sharing, works with a standard FedAvg backbone for the shared neural components, and allows the number of aggregated control embeddings to adapt to the system's missingness complexity — properties relevant to deployments over heterogeneous, privacy-constrained device fleets.

Future Directions

  • Scaling the theoretical analysis beyond the Lipschitz argument. The paper notes that L_ds constrains modality discrepancies in a shared embedding space but does not constrain the model's global behavior, leaving residual deviation governed by μ; tightening this remains open.
  • Extending beyond the two evaluated datasets. Results cover PTBXL (12 modalities) and Sleep-EDF (5 modalities); behavior on domains without a multimodal foundation model (e.g., healthcare) is argued but evaluated only on these benchmarks.
  • Understanding the aggregation of control pools further. The paper's ablation of the control pool and the visualization of global control embeddings over 500 iterations leave open questions about how cluster counts should evolve and how best to align profiles across very heterogeneous clients.
  • Handling mismatched train/test missing statistics more systematically. Table 2 shows PEPSY is robust when client and server missing statistics differ, but the paper does not report a mechanism designed specifically to anticipate mismatch; how to explicitly model that distribution shift is a natural next step.

Target Audience

Researchers and practitioners working on federated learning, multimodal learning, and robustness to missing data — particularly those building distributed systems over heterogeneous, incomplete, privacy-sensitive sensor data. The paper assumes familiarity with federated averaging, contrastive representation learning, and non-parametric clustering, making it most useful to readers at an intermediate-to-advanced level.

Authors’ abstract

Multimodal federated learning in real-world settings often encounters incomplete and heterogeneous data across clients. This results in misaligned local feature representations that limit the effectiveness of model aggregation. Unlike prior work that assumes either differing modality sets without missing input features or a shared modality set with missing features across clients, we consider a more general and realistic setting where each client observes a different subset of modalities and might also have missing input features within each modality. To address the resulting misalignment in learned representations, we propose a new federated learning framework featuring locally adaptive representations based on learnable client-side embedding controls that encode each client's data-missing patterns. These embeddings serve as reconfiguration signals that align the globally aggregated representation with each client's local context, enabling more effective use of shared information. Furthermore, the embedding controls can be algorithmically aggregated across clients with similar data-missing patterns to enhance the robustness of reconfiguration signals in adapting the global representation. Empirical results on multiple federated multimodal benchmarks with diverse data-missing patterns across clients demonstrate the efficacy of the proposed method, achieving up to 36.45\% performance improvement under severe data incompleteness. The method is also supported by a theoretical analysis with an explicit performance bound that matches our empirical observations. Our source codes are provided at https://github.com/nmduonggg/PEPSY

Read the original paper