Research
Explainable Multimodal Regression via Information Decomposition
Overview Research area: Multimodal machine learning, information-theoretic representation learning, and explainable AI (XAI) for continuous prediction. Technical level: Advanced. The paper builds on P

- arXiv
- 2512.22102
- Published
- 2025-12-26
- Authors
- Zhaozhao Ma, Shujian Yu
AI summary
Overview
- Research area: Multimodal machine learning, information-theoretic representation learning, and explainable AI (XAI) for continuous prediction.
- Technical level: Advanced. The paper builds on Partial Information Decomposition (PID), mutual information, Cauchy–Schwarz divergence, and the Shapiro–Wilk normality test.
- Scope: The paper proposes PIDReg, a framework that makes PID tractable for continuous, high-dimensional multimodal regression by enforcing joint Gaussianity of latent representations and the transformed target, yielding modality-level explanations alongside predictions.
What This Paper Is About
Multimodal regression models combine heterogeneous inputs (for example, audio, video, text, or imaging) to predict a continuous value, but they generally offer no principled way to say how much each modality contributes or whether modalities interact. The authors adapt Partial Information Decomposition, which splits mutual information into unique, redundant, and synergistic components, to the regression setting. Their goal is a model that both predicts accurately and explains, at the modality level, which inputs and cross-modal interactions drive the prediction.
Key Contributions
- A PID-based multimodal regression framework (PIDReg) that reveals the contributions of individual modalities and their higher-order interactions to the continuous output, rather than offering only post-hoc, instance-level explanations.
- An analytically tractable optimization scheme for PID on continuous high-dimensional variables. The authors resolve the underdetermined PID system by assuming the joint distribution of learned latent representations and the inverse-normal-transformed target is multivariate Gaussian, plus a closed-form conditional independence regularizer (via Cauchy–Schwarz divergence) that promotes unique information in each modality-specific encoder.
- Gaussianity enforcement mechanisms, including a Cauchy–Schwarz divergence penalty toward an isotropic Gaussian, a rank-based outlier-aware inverse normal transformation of the target, and a joint-Gaussianity regularizer built on whitening, vectorization, and the Shapiro–Wilk test.
- Extensive empirical validation on six real-world datasets spanning healthcare, physics, affective computing, and robotics, covering univariate and multivariate prediction, compared against six state-of-the-art methods, plus a synthetic study where redundancy, uniqueness, and synergy are controllable.
Main Findings
- Synthetic data recovery: PIDReg accurately estimates the relative strengths of the underlying generative factors. The paper reports that the estimated S/R ratio follows a monotonic trend with respect to the true w_s/w_r ratio, and that the estimated unique information U is near zero when the corresponding true weight w_u is approximately zero.
- CT Slices (medical imaging): PIDReg achieves the lowest RMSE (0.626) and highest correlation (1.000), versus MIB (1.801 / 0.997), MoNIG (1.490 / 0.996), MEIB (1.258 / 0.999), and DER (0.847 / 1.000).
- Superconductivity (physics): PIDReg achieves the lowest RMSE (10.37) and highest correlation (0.952), versus MIB (15.18 / 0.907), MoNIG (14.59 / 0.913), MEIB (14.04 / 0.917), and DER (12.37 / 0.936).
- CMU-MOSI and CMU-MOSEI (sentiment, three modalities: audio A, text T, vision V): With Audio&Text (MOSI: Acc7 32.0, Acc2 80.0, F1 79.7, MAE 0.938, Corr 0.662) and Visual&Text (MOSI: Acc7 37.2, Acc2 80.8, F1 80.9, MAE 0.947, Corr 0.664), PIDReg consistently outperforms all baselines. With Audio&Visual, the paper states PIDReg achieves second-best performance and notes that all competing methods exhibit a performance drop (MOSI Audio&Visual PIDReg: Acc7 16.4, Acc2 52.3, F1 51.8, MAE 1.400, Corr 0.149).
- Vision&Touch (robotics, multivariate prediction of a 4-dimensional (x, y, z, yaw) subset): Measured by MSE (×10⁻⁴) and RV coefficient Corr*, PIDReg reaches MSE 1.53 and Corr* 0.98, compared with MIB (3.00 / 0.97), MEIB (6.19 / 0.96), CoMM (1.34 / 0.98), MoNIG (3408 / 0.82), and DER (498 / 0.85).
- Brain age prediction on REST-meta-MDD (T1-weighted sMRI and rs-fMRI): PIDReg achieves the lowest MAE (6.29) and highest correlation (0.75), versus MIB (6.75 / 0.64), MoNIG (8.70 / 0.59), MEIB (7.83 / 0.65), CoMM (9.46 / 0.27), and DER (9.96 / 0.54). Results are averaged over nine experiment groups, each training on 15 randomly selected hospitals and testing on the remaining 2.
- Clinical consistency: The predicted age difference (PAD) histogram is reported to be consistent with existing medical evidence, where conditions such as Alzheimer's disease and major depressive disorder are associated with accelerated brain aging and larger PAD, following standard linear bias correction.
- Interpretable modality selection: Because fusion weights are derived directly from PID components and constructed as a linear combination, the framework supports dynamic modality selection, suppressing unreliable modalities and enabling informed modality selection for efficient inference.
Methodology in Plain English
Two modality-specific stochastic encoders map each input modality into an embedding. An adaptive linear–noise information bottleneck then regulates each embedding: a trainable scalar λ_m interpolates between the real embedding and Gaussian noise matched to its own batch mean and covariance, so that λ_m near 1 preserves information and λ_m near 0 suppresses it, all learned end-to-end without reparameterization.
The two bottleneck outputs are combined into a fused representation as a weighted sum of the first embedding, the second embedding, and their element-wise (Hadamard) product, which stands in for synergy. The fusion weights are not free parameters; they are computed from PID components (unique information for each modality, synergy, and redundancy), with redundancy split between the two modalities by a Bernoulli(0.5) switch so neither modality is systematically favored.
To compute those PID terms, the authors need mutual information expressions that are otherwise intractable for continuous high-dimensional variables. They adopt union information to close the underdetermined PID system, then restrict the optimization to Gaussian joint distributions of the latent representations and the transformed target, which yields a closed-form objective solvable with projected gradient descent. Three regularizers keep this assumption viable: a Cauchy–Schwarz divergence penalty pulling each latent marginal toward an isotropic Gaussian, a joint-Gaussianity penalty based on whitening, vectorization, and the Shapiro–Wilk statistic, and a conditional mutual information penalty (estimated in closed form with Cauchy–Schwarz divergence and Gram matrices) that discourages each encoder from encoding information belonging to the other modality.
Training uses a two-stage strategy: first everything is updated end-to-end until the fusion weights stabilize, then the fusion weights are frozen and the encoders and predictor are refined with stage-specific learning rates.
Why This Matters
- Impact on research: The paper shows that PID, long underdetermined for continuous variables, can be made usable inside an end-to-end deep network. This gives multimodal learning an intrinsic, modality-level explanation mechanism rather than a post-hoc one, addressing concerns the authors raise about the faithfulness of post-hoc explanations.
- Healthcare diagnostics: The brain age case study on REST-meta-MDD (848 MDD patients, 794 healthy controls, 17 hospitals) demonstrates use in neuroimaging-based biomarkers, where knowing whether structural MRI or functional MRI drives the prediction matters clinically.
- Medical imaging and clinical scoring: The CT Slice task (53,500 slices, 74 patients) predicts axial position from bone-structure and air-inclusion histograms, illustrating use where two imaging-derived feature sets must be weighed against each other.
- Affective computing and sentiment analysis: On CMU-MOSI (2,199 labels) and CMU-MOSEI (23,454 labels), the framework answers questions such as whether audio or text contributes more, and whether audio–visual pairs are synergistic or largely redundant.
- Physics and materials discovery: On Superconductivity (21,263 samples), it predicts critical temperature from chemical-property and chemical-formula modalities, where understanding which representation carries the signal guides further data collection.
- Robotics: On Vision&Touch (150 trajectories, a 7-DoF Franka Emika Panda robot), it links visual and force/torque signals to end-effector state, supporting sensor-suite decisions on robots.
- Industry relevance: Any pipeline that fuses several heterogeneous data sources for a continuous target, such as clinical risk scores, sensor fusion, or multimodal user modeling, can use the reported fusion weights to prune expensive modalities at inference time, which is directly relevant to deployment cost and latency.
Future Directions
- Scaling beyond two modalities: The paper states that the mechanism can be naturally extended to more than two modalities and discusses this in its appendix, but only two-modality settings are evaluated in the main text.
- Relaxing the Gaussian assumption: The main framework enforces joint Gaussianity of latents; the authors report results for non-Gaussian latents in an appendix, leaving broader distributional robustness an open question.
- Modality selection and efficient inference: The paper frames informed modality selection for efficient inference as a benefit; a systematic study of how aggressively modalities can be dropped at what accuracy cost is a natural next step.
- Broader validation: All results come from the six reported datasets; whether the PID estimates remain faithful on tasks with different noise structures, more than three modalities, or non-stationary data is not established.
Target Audience
Researchers and practitioners in multimodal learning, information-theoretic representation learning, and explainable AI who need modality-level rather than instance-level explanations. It is also relevant to applied scientists in medical imaging, neuroimaging, affective computing, materials physics, and robotics who fuse heterogeneous continuous signals and need to know which source actually carries predictive information. Readers should be comfortable with mutual information, partial information decomposition, and kernel-based divergence estimation, since the derivations are dense and the truncation of the provided content means the appendices, full algorithm listing, ablation studies, and the paper's concluding discussion are not available here.
Authors’ abstract
Multimodal regression aims to predict a continuous target from heterogeneous input sources and typically relies on fusion strategies such as early or late fusion. However, existing methods lack principled tools to disentangle and quantify the individual contributions of each modality and their interactions, limiting the interpretability of multimodal fusion. We propose a novel multimodal regression framework grounded in Partial Information Decomposition (PID), which decomposes modality-specific representations into unique, redundant, and synergistic components. The basic PID framework is inherently underdetermined. To resolve this, we introduce inductive bias by enforcing Gaussianity in the joint distribution of latent representations and the transformed response variable (after inverse normal transformation), thereby enabling analytical computation of the PID terms. Additionally, we derive a closed-form conditional independence regularizer to promote the isolation of unique information within each modality. Experiments on six real-world datasets, including a case study on large-scale brain age prediction from multimodal neuroimaging data, demonstrate that our framework outperforms state-of-the-art methods in both predictive accuracy and interpretability, while also enabling informed modality selection for efficient inference. Implementation is available at https://github.com/zhaozhaoma/PIDReg.