Skip to content
AI.info

Research

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Overview Research area: Multimodal machine learning, specifically robustness to missing modalities at deployment time, analyzed through the geometry of neural network parameter subspaces (Grassmannian

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness
arXiv
2610.04792
Published
2026-10-03
Authors
Songyuan Sui, Zhen Tan, Mohan Zhang, Rana Muhammad Shahroz Khan, Xia Hu, Tianlong Chen

AI summary

Overview

Research area: Multimodal machine learning, specifically robustness to missing modalities at deployment time, analyzed through the geometry of neural network parameter subspaces (Grassmannian manifolds, principal angles).

Technical level: Intermediate to Advanced. The empirical findings and practical method are accessible, but the theoretical framing relies on singular value decomposition, principal angles, Grassmannian geodesics, and machine unlearning terminology.

Scope: The paper diagnoses why full-modality-trained multimodal models degrade when a modality is absent at inference, attributes the degradation to principal-subspace rotations in cross-modal interaction layers, and proposes a lightweight post-training parameter edit called Geodesic Unlearning (GU) to correct it.

What This Paper Is About

Multimodal models are trained with all modalities present, but in real deployment a modality can disappear because of sensor failure, incomplete data collection, privacy constraints, or system limitations. The authors show that in these situations a model trained on text plus image can actually score worse than a model trained on text alone. The paper's goal is to explain this failure geometrically and to fix it with a targeted, low-cost edit to a single layer of an already-trained model.

Key Contributions

  1. A systematic characterization of deployment-time degradation. The authors document, across three major multimodal architecture families and multiple datasets, that full-modality inference performs best, unimodal training performs next, and full-modality training evaluated with a missing modality performs worst. This ordering holds in 26 of 27 architecture–dataset settings for the three-way ordering, and in all 27 settings for the comparison between unimodal training and missing-modality inference.

  2. A geometric account of the failure. Parameter-level analysis shows that multimodal training induces structured rotations of principal input subspaces, measured by principal angles between matched multimodal and unimodal models, especially in cross-modal interaction layers. Representation-level analysis shows larger task-aware harm (negative signed task margin) on misclassified samples at intermediate multimodal interaction stages under missing-modality inference, with the gap often persisting toward the classifier.

  3. Geodesic Unlearning (GU), a post-training subspace-editing method. GU selects one target layer using an angle–harm criterion, treats the selected principal subspace as a point on the Grassmannian, rotates its input-side basis toward a matched unimodal reference along a geodesic path, and reconstructs the layer while retaining the principal coefficients, output factor, and residual.

  4. A theoretical guarantee and broad empirical validation. The authors prove that the geodesic correction minimizes the distance to the reference within a fixed subspace-distance budget and that the path advances all principal angles proportionally. Experiments span fusion models, CLIP-style two-tower models, and BLIP-2 vision-language models.

Main Findings

  • A counterintuitive deployment failure: The ordering Perf(TI-TI) > Perf(T-T) > Perf(TI-T) holds in 26 of the 27 settings examined (TI-TI = train and test with text + image; T-T = train and test with text only; TI-T = train with text + image, test with text only). Perf(T-T) > Perf(TI-T) holds in all 27 settings. The single non-strict case is strong-encoder Fusion Cross-Attention on Hateful Memes, where TI-TI = T-T = 66.93 > TI-T = 66.85.

  • Size of the gap: Across the 27 settings, the missing-modality gap averages 4.74 F1 points but ranges from 0.08 to 20.29. Architectures with more explicit cross-modal interaction, such as cross-attention fusion, tend to show larger gaps than weakly coupled designs such as gated fusion or mean-pooled dual encoders.

  • Subspace rotation tracks degradation: After within-family z-score normalization, the mean principal angle of interaction-layer subspaces shows a strong positive association with the missing-modality gap (r = 0.91, R² = 0.82) over all 27 settings. Raw within-family measurements show consistent positive associations: fusion models (r = 0.93, R² = 0.86), stronger-encoder fusion models (r = 0.88, R² = 0.78), and CLIP-style two-tower models (r = 0.86, R² = 0.74). The BLIP-2 trend is positive but is described as descriptive only, because it contains three points.

  • Layer-level localization: On SNLI-VE, Fusion Cross-Attention exhibits a gap of 20.29 F1 points, with mean input/output principal angles of 16.00°/16.57° at attn_v_from_t_in_proj, versus 1.95°/0.00° at the final classifier.

  • Representation harm concentrates at interaction stages: Using a centroid-margin diagnostic, the gap between final-error and final-correct task-aware harm is already evident at cross-modal interaction stages before the final decision layer, and classifier-stage harm in several settings is interpreted as downstream persistence rather than a classifier-only effect. Bars in the corresponding figure are normalized within each dataset and model setting.

  • GU improves missing-image F1: Under TI-T evaluation on Hateful Memes, IU-XRay, and SNLI-VE (three-run means), GU achieves 60.88, 81.41, and 74.41 respectively, compared with null-token baselines of 57.63, 80.00, and 66.35; SMIL at 56.60, 79.21, and 72.20; Flex-MoE at 47.58, 71.02, and 70.74; DyMo at 48.82, 79.19, and 61.30; and layer interpolation at 58.50, 80.14, and 67.29. Gains over the strongest baseline are +2.38, +1.27, and +2.21 F1 points.

  • Baseline behavior differs by dataset: SMIL and Flex-MoE are competitive on SNLI-VE but do not consistently beat the null-token baseline across datasets, while DyMo degrades substantially in this strict bimodal setting where one of two modalities is entirely absent.

  • Robustness is gained without sacrificing full-modality accuracy: On SNLI-VE, GU improves TI-T by 3.49 points for Fusion while changing TI-TI by −0.09; by 1.69 for stronger-encoder Fusion with TI-TI change of −0.34; by 1.22 for CLIP with TI-TI change of −0.64; and by 0.01 for the VLM with a TI-TI change of +0.01. Comparable interpolation values are 0.98, 0.96, 1.04, and 0.01 for TI-T, with TI-TI changes of −0.13, −1.13, −1.21, and 0.00.

  • Gains are largest when dependencies are localized: GU is strongest for fusion models, where cross-modal dependencies sit in explicit fusion layers, and smaller for CLIP-style and BLIP architectures, where dependencies are more distributed.

  • Cost efficiency: The GU edit stage takes 230–858 seconds, reducing wall-clock time by 57–88% relative to post-hoc LoRA. Including matched reference training, GU uses 7–22% less wall-clock time and 16–62% less peak GPU memory than Dual CE across all three architectures. GU introduces no additional parameters or forward-pass FLOPs relative to the original TI model.

  • Ablation results: Varying the geodesic coefficient η from 0.25 to 1.0 causes only moderate changes in missing-modality performance and full-modality preservation; ranks from 2 to 32 yield similar TI-T performance with small TI-TI fluctuations. The target layer is the decisive factor, with attn_v_from_t_in producing the largest TI-T improvement while maintaining strong TI-TI performance.

  • Representation-level preservation: UMAP visualization on SNLI-VE shows GU and layer interpolation produce similar broad layouts; quantitative analysis in the appendix reports that both retain high similarity to the original TI-T representations while GU achieves larger error-margin gains and corrects more errors with a small increase in regressions.

Methodology in Plain English

The authors start by comparing three training-and-testing regimes for text-image tasks: train and test with both modalities (TI-TI), train and test with text only (T-T), and train with both but test with text only (TI-T). They measure both the missing-modality gap (TI-TI minus TI-T) and the unimodal robustness gap (T-T minus TI-T), using F1 as the task metric.

To find the cause, they take the weight matrix of a layer, compute its singular value decomposition, and look at the top-k input singular vectors — the dominant "reading directions" the layer uses. Because any orthonormal rotation of those vectors spans the same subspace, the mathematically meaningful object is the subspace itself, treated as a point on a Grassmannian. They compare the subspace of a multimodal-trained layer with that of a matched unimodal reference layer using principal angles.

Finding that the biggest rotations occur in cross-modal interaction layers, they then localize representational damage: for each intermediate stage they compute class centroids, a task direction, and a signed task margin, then compare average harm on samples the model gets wrong versus samples it gets right.

For the fix, GU picks one layer by ranking candidates on a development set with a score combining the z-scored mean principal angle and the z-scored harm gap. It splits the chosen weight matrix into a top-k principal part plus a residual, replaces the principal input basis with one rotated along the Grassmannian geodesic toward the unimodal reference basis, and rebuilds the layer keeping the original output factor, principal coefficients, and residual. A strength parameter η ∈ [0,1] controls how far along the path the edit travels. The test set is reserved for final evaluation only.

Why This Matters

Impact on research: The paper reframes missing-modality robustness as a model-side, parameter-level problem rather than a data-side recovery problem, and explicitly connects it to machine unlearning. It shows that multimodal training leaves harmful cross-modal dependencies in specific parameter subspaces, giving the field a geometric diagnostic — principal-angle deviation — that correlates with degradation. It also suggests that parameter editing can complement, rather than replace, training-time robustness objectives and modality-balancing methods such as OGM-GE, AGM, DnR, and MCR.

Real-world applications:

  • Clinical decision support where imaging may be unavailable but text records exist, relevant given the IU-XRay medical vision-language benchmark.
  • Multimodal content moderation and hateful-content detection where the image or text attachment is stripped by platform or privacy rules.
  • Visual entailment and image-text reasoning services deployed on heterogeneous devices with varying sensor availability.
  • Mobile or edge assistants where a camera stream drops out but a text query remains.

Industry relevance: The cost profile matters for production. GU is a one-time post-training edit with no added parameters and no extra inference FLOPs, and inference uses only the edited checkpoint without keeping the unimodal reference. That makes it attractive where retraining from scratch (as in Dual CE) or ongoing adaptation (as in post-hoc LoRA) is too expensive.

Future Directions

  • Extending GU beyond classification, including free-form multimodal generation, which the authors list as an important direction since current evaluation covers classification tasks and multiple-choice VQA.
  • Testing at larger scale: the authors note that due to resource constraints they do not run large-scale experiments on substantially larger models.
  • Addressing benchmark artifacts: concerns have been raised about rendered text in Hateful Memes and hypothesis-side artifacts in SNLI-VE. The authors report targeted preprocessing and complementary controls, and state that further evaluations on datasets without these artifacts reproduce both the degradation and GU's improvements.
  • Handling distributed dependencies: GU is weaker on CLIP-style and BLIP architectures, where harmful dependencies are more spread out than in a single fusion layer, leaving open how to correct non-localized damage.
  • Choosing which reference model to realign toward, and whether the same geometric correction transfers to settings with more than two modalities or partially missing samples, remains unaddressed in the reported content.

Target Audience

Researchers and engineers working on multimodal learning, robustness, and deployment reliability; practitioners who need to keep multimodal systems working when an input stream drops; and readers interested in geometric approaches to neural network analysis, parameter-efficient adaptation, or machine unlearning. It will also interest those studying modality imbalance and fusion architecture design, since the results indicate that stronger, more explicit cross-modal interaction can come with larger missing-modality gaps.

Authors’ abstract

Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.

Read the original paper