Research
Handling Missing Modalities in Multimodal Survival Prediction for Non-Small Cell Lung Cancer
Overview Research area: Multimodal deep learning for medical prognosis — specifically overall survival prediction in non-small cell lung cancer (NSCLC) using computed tomography (CT), whole-slide hist

- arXiv
- 2601.10386
- Published
- 2026-01-15
- Authors
- Filippo Ruffini, Camillo Maria Caruso, Claudia Tacconi, Lorenzo Nibid, Francesca Miccolis, Marta Lovino, Carlo Greco, Edy Ippolito, Michele Fiore, Alessio Cortellini, Bruno Beomonte Zobel, Giuseppe Perrone, Bruno Vincenzi, Claudio Marrocco, Alessandro Bria, Elisa Ficarra, Sara Ramella, Valerio Guarrasi, Paolo Soda
AI summary
Overview
Research area: Multimodal deep learning for medical prognosis — specifically overall survival prediction in non-small cell lung cancer (NSCLC) using computed tomography (CT), whole-slide histopathology images (WSI), and structured clinical (tabular) data, with a design that tolerates missing modalities.
Technical level: Intermediate (assumes familiarity with survival analysis metrics such as the concordance index, deep learning encoders, and the concept of foundation models).
Scope: The paper presents and evaluates a three-stage, missing-aware multimodal survival framework — foundation-model feature extraction, missing-aware unimodal encoding (NAIM+ODST), and intermediate fusion (ConcatODST) — on a retrospective cohort of 179 patients with unresectable stage II–IIIC NSCLC.
What This Paper Is About
Predicting how long a lung cancer patient will survive is more accurate when clinical records, CT scans, and biopsy slides are considered together, but real hospital data are rarely complete — in this cohort only 22.9% of patients had all three data types. Conventional multimodal models force researchers to either discard patients with incomplete data or artificially fill in the gaps, which shrinks the dataset and can bias results. The paper builds a framework that natively accepts whatever data a patient has, without imputation or exclusion, and shows that combining modalities in an intermediate fusion stage gives the best survival predictions.
Key Contributions
-
A missing-aware multimodal survival framework for NSCLC. The pipeline combines three stages — foundation-model feature extraction per modality (Step 1), a missing-aware transformer encoder with an ODST survival head (Step 2), and an intermediate concatenation-based fusion module, ConcatODST (Step 3) — that admits the full cohort by design rather than enforcing complete-case filtering or imputation.
-
Validation of foundation models as prognostic feature extractors in a small cohort. The work tests whether generalist representations from pretrained FMs (CT-FM for CT, a CLAM/MI-Zero/TITAN-based pipeline for WSI) can replace task-specific encoder training in a data-scarce oncological setting, and reports comparative analyses of CT foundation models (CT-CLIP, MERLIN, CT-FM) in supplementary material S.3.
-
A systematic comparison of unimodal, early, late, and intermediate fusion strategies. The study benchmarks machine-learning baselines (CPH, RSF, SGB) against deep-learning architectures (a plain NN, NAIM, and NAIM+ODST) under 5-fold cross-validation with fixed splits across all experiments.
-
A modality-importance analysis via test-time masking. Individual data streams are progressively masked at test time to observe the effect on survival prediction, providing an interpretable view of modality relevance, alongside risk-score stratification of disease progression and metastatic risk.
Main Findings
-
Intermediate fusion is the strongest strategy. The trimodal ConcatODST configuration reached a C-index of 74.42 ± 4.25, td-AUC of 80.74 ± 5.21, and Uno C-index of 69.01 ± 3.79, outperforming early fusion (73.26 ± 4.05 C-index for WSI+CT+Tab) and late fusion (71.50 ± 4.58 C-index for WSI+CT+Tab) in the trimodal setting.
-
Multimodal models beat unimodal baselines. The best unimodal scores were RSF on WSI (C-index 65.47 ± 1.15), SGB on CT (C-index 70.22 ± 5.05), and RSF on Tabular (C-index 67.47 ± 3.61) — all below the best multimodal configurations.
-
NAIM+ODST was the strongest deep-learning encoder across modalities. On CT it achieved the highest td-AUC (74.57 ± 6.42) of any model and a C-index of 68.75 ± 5.29; on Tabular it reached a C-index of 65.52 ± 4.45 and Uno of 63.56 ± 4.35, improving over standalone NAIM by 3.0 points in td-AUC (72.73 vs 69.68).
-
WSI was the hardest modality to model. With RSF as the best WSI model at C-index 65.47 ± 1.15, the plain NN baseline failed to generalize (Uno 45.21 ± 5.49, td-AUC 41.64 ± 5.95), and NAIM+ODST reached a C-index of 58.81 ± 4.04.
-
Representation-space analysis showed different levels of prognostic signal per modality. UMAP projections of FM embeddings showed no significant outcome-group separation for WSI (p = 0.2097) or tabular data (p = 0.2952), while CT embeddings showed statistically robust separation (p = 0.0421).
-
Missingness is structural and substantial. WSI was available for only 33.5% of the cohort (n = 60), CT was missing in 2.8% of cases, and at the modality level tabular data had 0% missingness. Only 22.9% of patients (n = 41) had a complete trimodal profile, so complete-case filtering would remove 77.1% of cases.
-
Tabular data had severe internal sparsity despite being nominally complete. EGFR was absent in 73.7% of patients, ALK in 73.2%, and MET in 96.6%.
-
Cohort differences aligned with survival. EGFR mutation status differed between groups (p = 0.033), with positive cases found exclusively among survivors; radiotherapy technique differed strongly (p < 0.001), with VMAT more common among survivors (60.3%) and 3D-CRT nearly three times more frequent among non-survivors (43.2% vs 16.2%); adjuvant immunotherapy was associated with improved survival (p = 0.010, 32.5% vs 15.2%). PFS events (p < 0.001, 78.8% of non-survivors vs 45.0% of survivors) and distant metastasis (p < 0.001, 64.6% vs 33.8%) were strongly linked to mortality.
-
Risk stratification was clinically meaningful. The learned risk scores produced statistically significant log-rank tests across all modality combinations according to the abstract, stratifying disease progression and metastatic risk.
Methodology in Plain English
The researchers used a retrospective cohort of 179 patients with unresectable stage IIA–IIIC NSCLC treated with radical chemoradiotherapy. They split modeling into three stages.
First, each data type is converted into a fixed numeric representation using pretrained foundation models rather than training encoders from scratch: a tabular preprocessing pipeline, CT-FM for CT volumes (producing a 512-dimensional token per volume), and a multi-stage WSI pipeline combining CLAM-based patch aggregation with MI-Zero vision encoders and TITAN (producing a 768-dimensional slide-level representation).
Second, each modality's representation passes through a modality-specific encoder built on NAIM, which uses missing-aware tokenization and masked self-attention so the network can structurally ignore unavailable inputs. A differentiable tree layer (ODST) replaces the linear survival head to capture non-linear interactions, producing a unimodal risk estimate.
Third, for the intermediate fusion model (ConcatODST), the unimodal encoders are frozen, their latent vectors are concatenated, and the combined representation passes through a multimodal module — a fully connected layer downscaling to one third of the input dimension followed by an ODST layer — and a final layer that outputs the patient-specific hazard.
All models were evaluated with 5-fold cross-validation using fixed splits, and reported metrics are mean ± standard error for the C-index, time-dependent AUC, and Uno's C-index. CPH and RSF handle missing tabular entries natively during fitting, while SGB required KNN-imputed data; for CT and WSI, patients lacking that modality were excluded only from the corresponding unimodal analysis. The framework itself never drops patients, so effective sample size varies by modality according to data coverage.
Why This Matters
Impact on research. The paper directly targets a tension that undermines most multimodal oncology studies: standard complete-case pipelines silently discard the majority of a real cohort, and the patients most likely to be excluded may be those with the most aggressive disease (since incomplete staging may itself reflect severity). Showing that a missing-aware architecture can retain the full cohort while still outperforming unimodal and alternative fusion strategies provides a template for how missingness can be treated as a first-class property of the model rather than a preprocessing nuisance. The finding that CT and tabular representations carry separable prognostic signal, while WSI does not in this cohort, also gives a concrete data point on whether generalist foundation-model features transfer to prognosis as opposed to diagnosis.
Real-world applications:
- Risk stratification at diagnosis for stage II–III NSCLC patients treated with chemoradiotherapy, to inform treatment intensity and follow-up scheduling.
- Identifying patients at elevated distant-metastasis risk who might benefit from intensified systemic therapy or closer surveillance.
- Supporting clinical decision-making in settings where diagnostic workups are routinely incomplete, since the model does not require all three data streams.
- Serving as a benchmark reference for research groups building multimodal survival models on small, real-world hospital cohorts.
Industry relevance. The framework is designed around foundation models used as frozen feature extractors, which reduces the labeled data needed and lowers the compute barrier for hospitals and small companies that cannot pretrain their own medical imaging encoders. A model that runs on whatever data are available also fits better into heterogeneous clinical IT environments, where CT, pathology, and structured records often live in separate systems with uneven coverage. The emphasis on interpretable modality masking further responds to the regulatory expectation that clinical AI systems explain which inputs drove a prediction.
Future Directions
- External and prospective validation. The study is retrospective and single-center with 179 patients; validating on external cohorts and prospectively is a logical next step to test generalizability.
- Improving WSI representations for prognosis. WSI embeddings showed no significant outcome-group separation (p = 0.2097) and were the weakest modality in unimodal prediction, raising the question of whether different aggregation strategies or prognosis-oriented pretraining would help.
- Characterizing the missingness mechanism. The paper notes that missing modalities may reflect disease severity itself; explicitly modeling why data are missing — rather than only that they are missing — could refine both bias assessment and predictions.
- Extending the framework beyond the three studied modalities. The architecture is described as modality-agnostic in structure, so incorporating genomic or additional imaging streams is an open avenue, as is deeper investigation of the modality-importance behavior summarized in the abstract.
Target Audience
This paper is most useful to researchers and practitioners working at the intersection of multimodal deep learning and clinical oncology — particularly those developing survival models on small, incomplete hospital datasets, those evaluating foundation models as feature extractors for prognosis rather than diagnosis, and clinical collaborators in radiation oncology, pathology, and radiology who need to judge whether such models can be trusted in realistic data conditions. It is also relevant to machine-learning engineers interested in missing-data-aware architectures that avoid imputation and complete-case filtering.
Authors’ abstract
Accurate survival prediction in Non-Small Cell Lung Cancer (NSCLC) requires integrating clinical, radiological, and histopathological data. Multimodal Deep Learning (MDL) can improve precision prognosis, but small cohorts and missing modalities limit its clinical applicability, as conventional approaches enforce complete case filtering or imputation. We present a missing-aware multimodal survival framework that combines Computed Tomography (CT), Whole-Slide Histopathology Images (WSI), and structured clinical variables for overall survival modeling in unresectable stage II-III NSCLC. The framework uses Foundation Models (FMs) for modality-specific feature extraction and a missing-aware encoding strategy that enables intermediate multimodal fusion under naturally incomplete modality profiles. By design, the architecture processes all available data without dropping patients during training or inference. Intermediate fusion outperforms unimodal baselines and both early and late fusion strategies, with the trimodal configuration reaching a C-index of 74.42. Modality-importance analyses show that the fusion model adapts its reliance on each data stream according to representation informativeness, shaped by the alignment between FM pretraining objectives and the survival task. The learned risk scores produce clinically meaningful stratification of disease progression and metastatic risk, with statistically significant log-rank tests across all modality combinations, supporting the translational relevance of the proposed framework.