Research
Fairness Evaluation of Risk Estimation Models for Lung Cancer Screening
Overview Research area: Machine learning fairness evaluation for medical imaging, specifically AI-based lung cancer risk estimation from low-dose CT (LDCT) screening. Technical level: Intermediate. Th

- arXiv
- 2512.22242
- Published
- 2025-12-23
- Authors
- Shaurya Gaur, Michel Vitale, Alessa Hering, Johan Kwisthout, Colin Jacobs, Lena Philipp, Fennie van der Graaf
AI summary
Overview
Research area: Machine learning fairness evaluation for medical imaging, specifically AI-based lung cancer risk estimation from low-dose CT (LDCT) screening.
Technical level: Intermediate. The paper assumes familiarity with AUROC, sensitivity/specificity, confidence intervals, and fairness terminology, but explains its models and framework accessibly.
Scope: A retrospective, two-stage fairness audit of three lung cancer risk estimation models (Sybil, Venkadesh21, PanCan2b) on a held-out NLST validation set of 5,911 scans, guided by the JustEFAB ethical framework.
What This Paper Is About
AI models can estimate lung cancer risk from LDCT scans and might ease the burden on a short-staffed radiology workforce, but the populations at high risk for lung cancer are diverse and it is unclear whether these models perform equally well across demographic groups. This study asks whether three widely used risk estimation models show performance disparities between subgroups (by sex, race, age, BMI, height, weight, and education), and then whether any such disparities count as unfair bias under an ethical framework, or are explained by genuine clinical differences in risk factors.
Key Contributions
-
A fairness evaluation of two deep learning lung cancer risk estimation models — Sybil (lung cancer risk over the next six years) and Venkadesh21 (pulmonary nodule malignancy risk) — plus the PanCan2b (Brock) logistic regression model recommended by the British Thoracic Society nodule management guideline, all assessed on the same held-out NLST validation set.
-
A two-stage methodology: Stage 1 measures subgroup performance disparities in AUROC, sensitivity, and specificity; Stage 2 tests whether those disparities persist after stratifying on the 10 clinical risk factors with the largest prevalence disparities between subgroups.
-
Operationalization of the JustEFAB framework's single negative-outcome criteria using separation-based fairness (per Barocas et al., 2023), focused on overestimation (false positives) and underestimation (false negatives) of lung cancer risk, and distinguishing fair clinical differences from ethically significant bias.
-
Publicly released analysis code, plus model-choice reasoning based on the fact that both deep learning models were trained on NLST data, allowing the authors to probe whether shared biases point to dataset causes and model-specific biases point to training-process causes.
Main Findings
-
Sybil differs significantly by sex: Sybil achieved an AUROC of 0.88 (95% CI: 0.86, 0.90) for women versus 0.81 (95% CI: 0.78, 0.84) for men, p < .001. At 90% specificity, sensitivity was 0.66 (0.60, 0.71) for women versus 0.53 (0.48, 0.58) for men (a 0.13 difference). At 90% sensitivity, specificity was 0.10 lower for men (0.46, 95% CI: 0.44, 0.48) than women (0.56, 95% CI: 0.54, 0.58).
-
Venkadesh21 shows a large racial sensitivity gap: At 90% specificity, sensitivity was 0.39 (95% CI: 0.23, 0.59) for Black participants versus 0.69 (95% CI: 0.65, 0.73) for White participants — a 0.30 difference. AUROC was non-significantly lower for Black participants (0.82, 95% CI: 0.74, 0.89) than White participants (0.89, 95% CI: 0.88, 0.91), p = 0.14.
-
BMI-related disparities: Sybil performed non-significantly better for participants with BMI of 25 kg/m² or above (AUROC 0.86, 95% CI: 0.84, 0.88) than below (AUROC 0.81, 95% CI: 0.78, 0.85), p = 0.05. PanCan2b showed a significant gap: AUROC 0.81 (0.79, 0.83) for BMI ≥ 25 versus 0.73 (0.70, 0.77) for BMI < 25, p = 0.001. All models had lower specificity for lower-BMI participants.
-
Height matters for Sybil: Sybil AUROC was 0.87 (95% CI: 0.85, 0.89) for participants 68 inches or shorter versus 0.80 (95% CI: 0.77, 0.83) for taller participants, p < 0.001, but there was little or no disparity by weight (p = 0.82).
-
PanCan2b and sex: PanCan2b showed slightly higher sensitivity for women than men but substantially lower specificity for women across thresholds, up to a 0.13 lower specificity at a 90% sensitivity threshold.
-
Age and education effects on false positives: All three models produced more false positives for participants older than 61 (0.07 lower specificity for Venkadesh21 and 0.12 lower for the other models at 90% sensitivity) and for participants who had not graduated high school.
-
Venkadesh21 at the Brock threshold: A 0.04 higher specificity for women at the Brock ILST threshold, but no other substantial sensitivity or specificity disparities and no significant AUROC disparity.
-
Biases attributed to unfairness: The sex disparity for Sybil and the racial sensitivity disparity for Venkadesh21 were not explained by available clinical confounders and therefore may be classified as unfair biases according to JustEFAB.
-
No significant disparity found for ethnicity/race in AUROC for Sybil or PanCan2b (p = 0.77 and p = 0.54 respectively), and no significant AUROC differences between education groups.
Methodology in Plain English
The authors took a single NLST validation set of 5,911 scans from 3,492 participants, constructed so that no scan appeared in any model's training data, and ran three risk models on it: Venkadesh21 (takes a 50 mm³ patch around a radiologist-annotated nodule), Sybil (takes a full LDCT scan and outputs six-year risk; the authors used the Year 1 score), and PanCan2b (a logistic regression using nodule, scan, and participant characteristics). For Venkadesh21 and PanCan2b, the maximum nodule score was used as the scan-level score.
Stage 1 split the cohort by demographic characteristics (median splits for age, height, weight; 25 kg/m² for BMI; high school graduation for education) and compared AUROC between subgroups with a two-tailed test from Hanley and McNeil (1982) at α = 0.05, and compared sensitivity and specificity at three thresholds: 90% overall sensitivity, 90% overall specificity, and the Brock "moderate risk" 6% threshold used in the ILST. Confidence intervals came from 1,000 bootstraps; a disparity was called substantial when subgroup confidence intervals did not intersect.
Stage 2 asked whether any disparity was fair or unfair. For each flagged disparity, the authors identified the 10 risk factors with the largest prevalence difference between the two subgroups, then re-measured model performance separately in scans where the factor was present and where it was absent. If the disparity survived in both subsets, the factor was not considered a confounder. If it disappeared and shrank in both, the factor was a potential confounder — but the authors note that under JustEFAB a factor linked to a demographic group through selection bias rather than biology, or unrelated to lung cancer risk, still counts as unfair.
Why This Matters
Research impact: The paper shows how to move beyond simply reporting a subgroup performance gap by layering an explicit ethical framework and confounder analysis on top, and by comparing models with shared training data (NLST) to distinguish data-driven from training-process-driven bias. It also demonstrates that a model can look acceptable on AUROC while hiding large sensitivity gaps at clinically used operating thresholds.
Real-world applications:
- National and regional lung cancer screening programs deciding whether to deploy AI risk estimation tools in radiology workflows.
- Radiologists and screening guideline committees applying nodule malignancy thresholds (such as the ILST 6% moderate-risk threshold and the British Thoracic Society's PanCan2b recommendation) to decide follow-up intervals.
- Regulators and health systems designing post-deployment monitoring of AI risk models across demographic subgroups.
- Developers auditing training data composition before a model reaches clinical use.
Industry relevance: Any organization building or procuring clinical decision support for screening needs subgroup-level evidence, not just aggregate AUROC. The paper also frames the practical stakes: overestimation burdens patients and clinicians with unnecessary workup and overdiagnosis, while underestimation can falsely label someone with cancer as low risk and deny them treatment, eroding trust in screening programs.
Future Directions
- Whether the observed Sybil sex disparity and Venkadesh21 racial sensitivity disparity can be reduced through retraining, reweighting, or calibration, and whether mitigation introduces trade-offs on other subgroups.
- How to expand fairness evaluation beyond NLST's demographically narrow validation set (93.4% White participants; only 188 Black participants in the validation set), since the small Black subgroup produced wide confidence intervals.
- How fairness findings should translate into practice when the framework's criteria conflict or when a clinically legitimate factor (such as age, a clinically established cancer risk factor) also produces disparate false positive rates.
- Whether the JustEFAB-based analysis should be extended to label annotation practices and screening selection criteria, which the authors explicitly placed out of scope.
Target Audience
Clinical AI researchers and machine learning fairness practitioners; radiologists and lung cancer screening program leads; medical AI regulators and health technology assessment groups; and developers of deep learning risk models for medical imaging who need a template for subgroup performance and confounder analysis. Readers without a medical imaging background can follow the framing but should expect familiarity with ROC-based metrics.
Authors’ abstract
Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95% CI: 0.86, 0.90) and men (0.81, 95% CI: 0.78, 0.84, p < .001). At 90% specificity, Venkadesh21 showed lower sensitivity for Black (0.39, 95% CI: 0.23, 0.59) than White participants (0.69, 95% CI: 0.65, 0.73). These differences were not explained by available clinical confounders and thus may be classified as unfair biases according to JustEFAB. Our findings highlight the importance of improving and monitoring model performance across underrepresented subgroups, and further research on algorithmic fairness, in lung cancer screening.