Skip to content
AI.info

Research

Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning

Overview Research area: Medical computer vision / multimodal deep learning for surgical oncology, specifically automated assessment of pancreatic cancer resectability from CT imaging combined with cli

Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning
arXiv
2607.13826
Published
2026-07-15
Authors
Vincent Ochs, Christoph Kuemmerli, Florentin Bieder, Julia Wolleb, Joel L. Lavanchy, Julia Ruppel, Jan Liechti, Stephanie Taha-Mehlitz, Christian Andreas Nebiker, Beat Mueller, Giuseppe Kito Fusai, Joerg-Matthias Pollok, Anas Taha, Philippe C. Cattin, Sebastian Staubli

AI summary

Overview

  • Research area: Medical computer vision / multimodal deep learning for surgical oncology, specifically automated assessment of pancreatic cancer resectability from CT imaging combined with clinical data.
  • Technical level: Intermediate. The clinical motivation is accessible, but the method involves Swin-UNETR backbones, multitask loss weighting, and nested cross-validation.
  • Scope: The paper introduces and evaluates a fully automated multimodal framework that classifies pancreatic ductal adenocarcinoma (PDAC) patients into the three NCCN resectability categories using 3D contrast-enhanced CT and 17 structured clinical variables.

What This Paper Is About

Deciding whether a pancreatic tumor can be surgically removed depends on how it touches major blood vessels around the pancreas, and the National Comprehensive Cancer Network (NCCN) defines three categories: upfront resectable, borderline resectable, and locally advanced. Radiologists and tumor boards frequently disagree on these categories, with reported inter-observer agreement often below 70% even under standardized criteria, leading to qualitative rather than reproducible judgments. This paper builds a deep learning system that reads the CT scan and the patient's clinical record together to assign one of the three NCCN categories automatically.

Key Contributions

  1. A clinically motivated multimodal framework that integrates 3D contrast-enhanced CT with 17 structured clinical variables for NCCN resectability prediction.
  2. Anatomy-guided auxiliary supervision: the shared encoder is trained to segment pancreas, tumor, and vascular structures, so it learns vessel-aware features, yet segmentation masks are required only during training and not at inference.
  3. A performance-adaptive multitask objective that dynamically shifts the balance between segmentation and classification losses based on the current tumor Dice score, acting like a curriculum that teaches anatomy first and classification later.
  4. A comprehensive evaluation including ablation studies (feature isolation, modality knockout, single-modality training, loss-weighting variants) and comparisons against an adapted transformer baseline (TAT) and a segmentation-based geometric baseline, plus external validation on an independent institution's cohort.

Main Findings

  • Internal cross-validation performance: In the 159-patient cohort (85 upfront resectable, 47 borderline resectable, 27 locally advanced), the method achieved a mean AUC of 0.86 ± 0.03, macro-F1 of 0.79 ± 0.02, and accuracy of 0.85 ± 0.03 using stratified nested 5-fold cross-validation.
  • External validation: On an independent 52-patient cohort from Kantonsspital Aarau (28 upfront resectable, 14 borderline resectable, 10 locally advanced), the model reached an AUC of 0.86, macro-F1 of 0.81, and accuracy of 0.87, supporting cross-institution generalization.
  • Outperforming baselines: The full multimodal model beat the adapted Texture-Aware Transformer (internal AUC 0.83 ± 0.03, macro-F1 0.77 ± 0.03, accuracy 0.81 ± 0.03) and the segmentation-based approach of Viviers et al. (internal AUC 0.79 ± 0.04, macro-F1 0.74 ± 0.03, accuracy 0.76 ± 0.04). The same ordering held on external validation.
  • Auxiliary segmentation quality: The Swin-UNETR decoder reached Dice scores of 0.82 ± 0.03 for pancreas, 0.71 ± 0.05 for tumor, and 0.67 ± 0.06 for major vessels, despite segmentation never being used at test time.
  • Both modalities carry complementary signal: Feature isolation gave AUC 0.83 ± 0.03 for frozen imaging features versus 0.75 ± 0.04 for frozen tabular features; single-modality retraining gave AUC 0.82 ± 0.03 for CT-only and 0.74 ± 0.04 for tabular-only, both below the full multimodal 0.86 ± 0.03.
  • Imaging is the harder modality to lose: In knockout tests, replacing the tabular embedding with its mean left performance at AUC 0.83 ± 0.03, while replacing the imaging embedding with its mean dropped performance to 0.77 ± 0.04.
  • Adaptive loss weighting helps: Fixed weighting schemes gave AUC 0.84 ± 0.03, 0.83 ± 0.03, and 0.80 ± 0.04 depending on the configuration, with mean tumor Dice of 0.49 ± 0.05, 0.52 ± 0.04, and 0.42 ± 0.06 respectively; the adaptive schedule achieved AUC 0.86 ± 0.03 and the highest tumor Dice of 0.56 ± 0.04.
  • Class weighting matters for the rarer categories: Unweighted cross-entropy gave AUC 0.83 ± 0.03 and macro-F1 0.75 ± 0.03, strict inverse-frequency weighting gave 0.85 ± 0.03 and 0.77 ± 0.03, and the smoothed weights (1, 2, 2.5) gave the best result at 0.86 ± 0.03 and 0.79 ± 0.02.
  • Per-class balance: F1-scores were 0.83 ± 0.04 for upfront resectable, 0.78 ± 0.06 for borderline resectable, and 0.70 ± 0.07 for locally advanced, showing the smallest class is the hardest.
  • Imputation did not appear to bias results: The external KSA Aarau cohort had complete clinical data for all model variables and required no imputation, yet performance was comparable to the training cohort where KNN-based imputation was applied.
  • Learned geometry resembles NCCN rules: A post-hoc analysis of tumor-vessel contact angles derived from predicted segmentations showed the three classes clustering in the expected regions, even though no angle-based supervision was used.

Methodology in Plain English

The researchers built a system that takes two inputs: a 3D CT scan of the abdomen and a table of routine clinical facts about the patient (things like age, BMI, comorbidity score, CA 19-9 and CEA tumor markers, diabetes status, nicotine and alcohol use, bilirubin, HbA1c, histopathological grade, and whether the patient had neoadjuvant therapy).

The CT scan is resampled to 1 mm spacing, cropped or padded to 160 × 160 × 160 voxels, intensity-normalized, and randomly augmented with flips, rotations, noise, and intensity shifts to make the model robust across scanners.

The imaging path uses a Swin-UNETR encoder-decoder, a transformer-based 3D network pretrained on the BTCV dataset. The encoder compresses the scan into a 256-dimensional feature vector, and the decoder is trained to segment 16 anatomical and pathological structures — pancreas, tumor, portal vein, splenic vein, superior mesenteric vein, superior mesenteric artery, celiac trunk, aorta, inferior vena cava, several hepatic and splenic arteries, the pancreatic duct, and the bile duct. These masks came from a multi-expert annotation workflow involving up to five physicians.

The clinical path passes the 17 variables through a small three-layer network to produce a 32-dimensional embedding. The two embeddings are concatenated into a 288-dimensional vector and fed to a classifier that outputs one of the three NCCN categories. The key trick is that gradients from both the segmentation and the classification task flow back into the same shared encoder, so the encoder is pushed to learn vessel-aware anatomy that is also useful for the final decision.

To keep training stable, the researchers made the loss weights depend on the current tumor Dice score: when segmentation is poor (Dice below 0.1) the segmentation loss gets weight 3.0, at mid-training (Dice 0.1 to 0.5) it drops to 1.5, and once Dice exceeds 0.5 it settles at 1.0, with the classification weight always set to the inverse. Because the classes are imbalanced, the classification loss uses smoothed inverse-frequency weights of 1.0, 2.0, and 2.5 for the three categories.

Evaluation used stratified nested 5-fold cross-validation: each outer fold held out 20% of the data untouched, while the remaining 80% was split into 64% training and 16% inner validation used only for model selection and early stopping. Models were trained in PyTorch with MONAI on an NVIDIA A100 (40GB) GPU, and inference used sliding-window processing with Gaussian blending. Segmentation outputs were used only for evaluation — the deployed classifier does not need them.

Why This Matters

Impact on research: This is, according to the authors, the first end-to-end integration of 3D vessel-aware CT features with structured clinical covariates specifically tailored to NCCN-defined resectability assessment. It shows that anatomical supervision can strengthen a classification model without imposing a segmentation burden at inference time, and it demonstrates that a model trained at two Basel centers transfers to an independent hospital with essentially unchanged performance.

Real-world applications:

  • Decision support in multidisciplinary tumor boards, where surgical, radiological, and oncological opinions must converge on a resectability category.
  • Standardizing CT interpretation across readers, addressing the reported inter-observer agreement below 70% under standardized criteria.
  • Triaging patients toward upfront surgery versus neoadjuvant (radio-)chemotherapy based on a reproducible category assignment.
  • A second-opinion or quality-assurance tool at institutions that lack specialized hepatopancreatobiliary expertise on site.

Industry relevance: The work is directly relevant to medical imaging AI vendors, hospital radiology and surgical workflow software, and clinical decision support platforms. The authors make the implementation publicly available at https://github.com/vincentochs/pancreas_resectability, and state that data and weights can be made available by the corresponding author upon reasonable request — a distribution model that matters for regulatory and translational pathways.

Future Directions

  • Multi-institutional evaluation on larger and more diverse cohorts, since the authors identify the moderate cohort size as a limitation and state that further validation on additional external datasets is needed.
  • A prospective multi-reader study comparing individual clinician performance against the model under standardized reading conditions, including interobserver agreement and comparison to multidisciplinary tumor board consensus, which the authors call an important next step.
  • Explainability analyses and formal feature-importance assessment to evaluate generalization and clinical robustness, which the authors flag as essential.
  • Incorporating additional imaging modalities such as multiphase CT or histopathology, and revisiting more expressive fusion operators like cross-attention — which did not improve performance and reduced training stability at the current cohort size but might help with more data.

Target Audience

This paper is most useful to hepatopancreatobiliary surgeons, radiologists, and oncologists interested in objective resectability assessment; to medical imaging and computer vision researchers working on multimodal fusion and anatomy-guided representation learning; and to clinical AI developers and regulators evaluating decision-support tools for surgical planning. Readers without a deep learning background can follow the clinical framing and results, but the methodological details assume familiarity with transformer-based 3D architectures and multitask training.

Authors’ abstract

Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability. We introduce a fully automated multimodal deep learning framework that jointly analyzes 3D contrast enhanced CT and structured clinical information to classify patients into the three National Comprehensive Cancer Network (NCCN) resectability categories (upfront resectable, borderline resectable, locally advanced). The approach uses a Swin-UNETR backbone to obtain anatomy aware image representations through auxiliary segmentation of pancreas, tumor, and vascular structures. These features are fused with a compact clinical embedding derived from 17 routinely collected variables and processed by a lightweight classification head. Model training is guided by a dynamic multitask objective that adapts the balance between segmentation and classification based on current tumor Dice performance, promoting feature representations that remain both anatomically informed and discriminative.

Read the original paper