Skip to content
AI.info

Research

Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions

Overview Research area: Explainable AI / auditing of deep learning models in medical imaging, specifically lung cancer risk prediction from low-dose CT (LDCT) scans, combining game-theoretic attributi

Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
arXiv
2602.02560
Published
2026-01-30
Authors
Bartlomiej Sobieski, Jakub Grzywaczewski, Karol Dobiczek, Mateusz Wójcik, Tomasz Bartczak, Patryk Szatkowski, Przemysław Bombiński, Matthew Tivnan, Przemyslaw Biecek

AI summary

Overview

Research area: Explainable AI / auditing of deep learning models in medical imaging, specifically lung cancer risk prediction from low-dose CT (LDCT) scans, combining game-theoretic attribution (n-Shapley values) with generative diffusion bridge modeling.

Technical level: Advanced. The paper assumes familiarity with Shapley values and interaction indices, diffusion models and SDEs, diffusion bridges, counterfactual explanations, and statistical testing (ANOVA, Tukey's HSD).

Scope: The paper proposes S(H)NAP, a model-agnostic interventional auditing framework that uses 3D diffusion bridge interventions on pulmonary nodules to causally explain — and expose failure modes in — Sybil, a deep learning model that predicts 6-year lung cancer risk from a single CT scan.

What This Paper Is About

Sybil is a frontier deep learning model that predicts 6-year lung cancer risk solely from a single CT scan, and it has been validated extensively in retrospective and prospective studies. However, all existing evaluations are observational: they confirm that the model works but say nothing about why it makes individual decisions or when it might fail. This paper builds a causal auditing framework that replaces or inserts pulmonary nodules using realistic 3D generative interventions, then measures how Sybil's risk score changes, in order to verify the model's reasoning against expert radiological judgment.

Key Contributions

  1. Novel synthetic intervention methodology: The authors introduce methods for intervening on pulmonary nodules in 3D LDCT using diffusion bridge modeling, grounded in the distribution-blending result of Verdú (2009) and the System-Embedded Diffusion Bridges (SDB) formulation. These enable both removal of nodules (replacing them with healthy tissue) and insertion of nodules with controlled properties.

  2. Two new attribution methods: SHNAP (SHapley Nodule Attribution Profiles), which decomposes Sybil's prediction into nodule main effects and pairwise interactions by removing every coalition of nodules, and SNAP (Substitutive Nodule Attribution Probing), which probes spatial sensitivity by inserting nodules of known malignancy at target coordinates. A generalized variant, gSHNAP, extends the same logic to arbitrary image regions.

  3. Empirical verification that Sybil behaves as a linear model with pairwise interactions (LMPI) over pulmonary nodules: The authors formalize this as Hypothesis 1 and test it with R² (unweighted local fidelity) computed over the feature lattice of nodule coalitions.

  4. Discovery of critical misalignments: The audit reveals a radial sensitivity bias (reduced sensitivity near the pleura), dangerous sensitivity to clinically unjustified artifacts (metal snaps, ECG electrodes, the mandible, thyroid goiters), and cases of "right for wrong reasons" where correct predictions rest on flawed reasoning.

Main Findings

  • Sybil is effectively a linear model with pairwise interactions over nodules. SHNAP was computed on the entire LUNA25 test split and iLDCT. Main effects alone sufficed for a perfect fit in the majority of cases, evidenced by a collapsed interquartile range at R² ≈ 1; adding pairwise interactions almost entirely eliminated the outlier tail. Remaining failures typically corresponded to rare, anomalously large nodules where SDB struggles with reconstruction.

  • Nodule removal is indistinguishable from real tissue. In a blinded evaluation, two board-certified radiologists assessed 40 3D cubes from NLST (20 real healthy tissue, 20 synthetic removal) in a binary classification task. Their performance was statistically indistinguishable from random guessing (exact binomial test, point estimate 0.57), and a response-bias analysis revealed a significant tendency to label samples as "real."

  • Insertion has an optimal alignment parameter of τ = 0.3. Radiologists evaluated 110 pairs of scans (original vs. inserted) across τ ∈ {0, 0.1, …, 1.0}, judging structural properties, malignancy characteristics, and background alignment. The transition runs from artifact-heavy naive insertion (τ → 0) to excessive deviation (τ → 1), with τ = 0.3 preserving source properties without visual artifacts.

  • SHNAP attributions are stable. Across 5 independent runs, the density of attribution standard deviations concentrated heavily around zero. Naive (non-generative) perturbations produced unstable attributions, highlighting the necessity of the generative intervention approach.

  • Risk attribution differs sharply between benign and malignant cases. Using the Relative Nodule Contribution (RNC), defined with a decision threshold set to enforce ≥ 95% sensitivity, the LUNA25 test split showed a highly right-skewed RNC distribution for benign cases (a long positive tail signaling potential flaws where benign nodules erroneously drive high risk), while confirmed cancers showed a bimodal distribution with a significant mode near 0, indicating Sybil frequently ignores known lesions in favor of background context. On iLDCT, with higher severity, focus shifts more toward nodules, but with a larger tail of erroneous high attributions for benign findings.

  • Documented false positives and false negatives. In one false positive, nodules alone drove 60% of the risk, with Sybil focusing almost entirely on a large dense lesion in the upper right lung while ignoring two smaller left-lung nodules. In another, 78% of risk stemmed from nodules, dominated by a single lesion with a dense center and ground-glass opacity; three-year longitudinal follow-up revealed the model's focus shifting to a different nodule, exposing inconsistency. A malignant case was severely underestimated with a total nodule contribution of merely 11%, with both subpleural nodules failing to receive significant weight.

  • Correct predictions do not imply correct reasoning. In one malignant case, Sybil assigned negative attribution to the actual nodules, treating them as evidence against cancer; only their interaction negated this effect. The correct high-risk prediction was driven almost entirely by background features.

  • Attention often lands on clinically unjustified artifacts. In one example, 5 attention-based regions were found, 2 of them outside the patient's body: the dominant region pointed to an artifact resembling metal snaps used to close a hospital gown, and another pointed to a cross-section of the patient's chin (mandible). In another benign case, 50% of the predicted risk stemmed from the joint influence of two symmetric objects outside the body, revealing ECG electrodes attached to the chest skin.

  • Influential regions are sparse. gSHNAP explanations for random regions within the lung volume disjoint from nodules and attention-based regions showed importance highly concentrated around zero, confirming the model is not simply reacting to arbitrary changes.

  • SNAP reveals local and global spatial biases. A high-resolution map of over 5,000 insertions of a malignant nodule in a single patient showed variable sensitivity: a site at the lung base triggered a strong response, a site near the pleura showed a weaker response, and a third site showed complete failure to detect the nodule. Aggregating 240 patient-nodule combinations (20 nodules × 12 scans, ≈900 insertions per combination, totaling ≈200,000 samples), a two-way ANOVA found significant main effects for patient identity (p < 0.001) and lobe class (p < 0.001), with a non-significant interaction (p ≈ 1.0), indicating lobar bias is a global characteristic of Sybil.

  • Lobar biases align with clinical standards. A post-hoc Tukey's HSD test showed attribution in the left and right upper lobes was significantly higher than in the middle/lower lobes (p ≤ 0.009), consistent with clinical models such as PanCan and Mayo that identify upper lobe location as a significant malignancy predictor. The right middle lobe was indistinguishable from the lower lobes (p-value ≈ 1.0), and Sybil correctly ignored laterality.

  • A radial sensitivity bias exists. A linear regression predicting attribution from distance-to-pleura yielded a significant positive coefficient (p < 0.001), meaning sensitivity drops near the lung boundary. Distance alone explained little variance (R² = 0.071), but adding interactions with nodule identity raised this to R² = 0.455, confirming malignant nodules are progressively attenuated near the boundary while benign ones remain robust. The authors hypothesize this stems from zero-padding in 3D convolutions.

  • The background term correlates weakly with patient age. A preliminary mixed-effects regression linked the baseline term μ_x to patient age, showing a positive trend on LUNA25 (p = 0.05) and iLDCT (p = 0.027).

Methodology in Plain English

The authors start from a clinical consensus: pulmonary nodules are the primary predictive biomarkers for lung cancer. They hypothesize that Sybil's decision function can be approximated by a simple equation containing a patient-specific background term plus additive contributions from individual nodules and pairwise interactions between nodules. Testing that equation requires counterfactual scans — the same patient with nodules selectively added or removed — which do not exist in paired form.

To generate them, the researchers train a 3D diffusion bridge model (a discrete-time, 1000-step Schrödinger Bridge variant of SDB trained on randomly sampled 64³ cubes, with training masks generated procedurally via metaballs) on NLST data. Diffusion bridges let them edit only a masked region while leaving the rest of the scan untouched. For nodule removal, they mask the nodule and let the model fill it with healthy tissue, justified by a theorem showing that the distinguishing information between the original and counterfactual distributions vanishes over the forward diffusion process. For insertion, they copy a nodule from another patient into a target location and tune a timestep parameter τ that balances preserving the source nodule's properties against blending it into the new anatomy.

With these interventions they build SHNAP: they remove every possible coalition of nodules, query Sybil on each resulting scan, and fit n-Shapley values (with n = 2, using an interventional value function) to recover main effects and pairwise interactions, measuring fit quality with an R² over the feature lattice. They also build SNAP: they insert a nodule of known malignancy at many target coordinates and record the change in Sybil's base-hazard logit as the attribution score. A generalized version, gSHNAP, replaces nodule masks with Sybil's own binarized attention regions to probe non-nodule areas. Finally, they validate the interventions with two board-certified radiologists in blinded studies, and analyze the resulting attributions with standard statistical tests (exact binomial tests, two-way ANOVA, Tukey's HSD, linear and mixed-effects regression).

Datasets used: NLST (approximately 28,000 training and 6,000 test scans) for training the generative model; LUNA25 (4,069 scans, with biopsy-confirmed malignancy or 2-year stability verification for benign cases) using naive spherical masks from nodule coordinates; and iLDCT (243 scans), an internal out-of-distribution testbed with higher prevalence of severe cases and precise expert radiologist annotations. Lung and lobe segmentations came from the improved TotalSegmentator. Nodule removal and insertion were performed with 100 NFE.

Why This Matters

Impact on research: The paper shifts medical AI evaluation from observational validation ("the model works") to interventional, causal auditing ("why and when it fails"). It provides a reusable, model-agnostic framework that relies only on input-output pairs, and it explicitly connects generative counterfactual methods to the theoretical distribution-blending result of Verdú (2009), a justification the authors note is rarely formulated explicitly despite being implicitly relied upon in image editing work.

Real-world applications:

  • Pre-deployment safety review of lung cancer screening AI, checking for reliance on artifacts or spatial blind spots before clinical use.
  • Radiologist trust calibration, by showing which nodule morphologies and locations the model weights correctly versus ignores.
  • Scanner- and protocol-level quality control, since the identified failure modes (metal snaps, ECG electrodes, the chin, thyroid goiters) are acquisition artifacts that could be removed or flagged.
  • Regulatory submission evidence, offering a structured causal audit trail rather than only area-under-the-curve style observational metrics.

Industry relevance: Because S(H)NAP is model-agnostic and depends only on input-output pairs, the authors state it can be applied to arbitrary systems, including proprietary commercial models such as Optellum. The finding that Sybil is essentially a linear model with pairwise interactions over nodules suggests that nodule-level explanations may be efficiently computable, and the discovery of spurious "hospital tag"-like shortcuts is directly relevant to vendors seeking clinical acceptance and regulatory clearance.

Future Directions

  • Overcoming reliance on partially synthetic data. The authors identify generative artifacts as a primary limitation, mitigated only by a blinded expert study, and note that provably robust counterfactuals remain an active frontier.
  • Improving reconstruction of anomalously large nodules. The residual poor-fit cases in the LMPI approximation correspond to rare, very large nodules where SDB struggles; the authors suggest this is addressable by training on larger volumes.
  • Disentangling global from local features. The preliminary finding that the background term μ_x tracks patient age suggests strong predictive power in global cues like bone density; the authors call for future work to separate these from localized pathology.
  • Explaining and, presumably, correcting structural biases. The hypothesized link between the radial sensitivity bias and zero-padding in 3D convolutions raises the question of whether architectural changes could remove a clinically critical blind spot, given that adenocarcinoma predominantly arises in the lung periphery.

Target Audience

This paper is most valuable to machine learning researchers working on explainable AI and counterfactual explanation methods; medical AI developers and auditors preparing models for clinical deployment; radiologists and clinical informatics specialists interested in the reasoning behavior of lung cancer screening models; and regulatory or safety reviewers who need causal, intervention-based evidence rather than observational performance metrics. Readers should be comfortable with Shapley values, diffusion models, and statistical inference.

Authors’ abstract

Lung cancer remains the leading cause of cancer mortality, driving the development of automated screening tools to alleviate radiologist workload. Standing at the frontier of this effort is Sybil, a deep learning model capable of predicting future risk solely from computed tomography (CT) with high precision. However, despite extensive clinical validation, current assessments rely purely on observational metrics. This correlation-based approach overlooks the model's actual reasoning mechanism, necessitating a shift to causal verification to ensure robust decision-making before clinical deployment. We propose S(H)NAP, a model-agnostic auditing framework that constructs generative interventional attributions validated by expert radiologists. By leveraging realistic 3D diffusion bridge modeling to systematically modify anatomical features, our approach isolates object-specific causal contributions to the risk score. Providing the first interventional audit of Sybil, we demonstrate that while the model often exhibits behavior akin to an expert radiologist, differentiating malignant pulmonary nodules from benign ones, it suffers from critical failure modes, including dangerous sensitivity to clinically unjustified artifacts and a distinct radial bias.

Read the original paper