Skip to content
AI.info

Research

On the Role of Calibration in Benchmarking Algorithmic Fairness for Skin Cancer Detection

Overview Research area: Algorithmic fairness benchmarking in medical AI, specifically skin cancer (melanoma) detection from dermoscopy images. Technical level: Intermediate. Readers should be comforta

On the Role of Calibration in Benchmarking Algorithmic Fairness for Skin Cancer Detection
arXiv
2511.07700
Published
2025-11-10
Authors
Brandon Dominique, Prudence Lam, Nicholas Kurtansky, Jochen Weber, Kivanc Kose, Veronica Rotemberg, Jennifer Dy

AI summary

Overview

Research area: Algorithmic fairness benchmarking in medical AI, specifically skin cancer (melanoma) detection from dermoscopy images.

Technical level: Intermediate. Readers should be comfortable with AUROC, the idea of model calibration, and basic statistical testing (p-values, confidence intervals), but the paper's framing of the problem is accessible to a general machine learning audience.

Scope: The paper audits the top three models from the ISIC 2020 melanoma classification challenge, plus a standard ERM baseline, on two dermoscopy datasets, adding calibration-based fairness analysis to the AUROC-based group fairness metrics that dominate prior benchmarking work.

What This Paper Is About

AI models can detect melanoma at expert level, but they perform unevenly across demographic groups such as sex, race and age, which blocks clinical adoption. Prior auditing efforts rely almost entirely on AUROC-based group fairness metrics, which measure how well a model ranks cases but say nothing about whether the probabilities it outputs are numerically trustworthy. This paper asks whether adding calibration as a complementary benchmark metric reveals subgroup problems that AUROC-only auditing misses.

Key Contributions

  1. Calibration as a complementary fairness benchmark. The paper pairs standard AUROC-based group fairness analysis with a calibration analysis, arguing that discrimination and calibration answer different clinical questions and must be audited together.

  2. An adaptive CUSUM calibration test applied to skin cancer detection. Rather than checking calibration only on a predefined list of subgroups, the authors use the score-based CUSUM test of Feng et al. (2023), which scans all subgroups present in the audit dataset in one pass, with Variable Importance (VI) plots to attribute detected miscalibration to specific features.

  3. A head-to-head audit of the top three ISIC 2020 challenge models plus an ERM baseline. The first-place ADAE ensemble, the 2nd Place and 3rd Place ensembles, and a ResNet-18 ERM baseline are compared on the ISIC 2020 Challenge dataset and on the external PROVE-AI dataset, with intersectional subgroup analysis.

  4. Public code release. All code is available at https://github.com/bdominique/testing_strong_calibration.

Main Findings

  • The top three challenge models beat the baseline on ISIC 2020. All three significantly outperformed ERM across every subgroup on that dataset (all p-values reported as 0.000 or low single-digit values), with overall AUROCs of 0.949 (ADAE), 0.949 (2nd Place), 0.953 (3rd Place) versus 0.866 for ERM.

  • The top three were nearly indistinguishable from each other on ISIC 2020. Comparisons among them were generally small and not statistically significant (p > 0.05), with only a few marginal subgroup results.

  • ADAE over-diagnoses. On ISIC 2020 at the 95% sensitivity threshold, ADAE produced the most false positives across subgroups (3157 overall, versus 2469 for 2nd Place, 2294 for 3rd Place and 5864 for ERM) and had the lowest specificity among the top three (71%, versus 77% and 79%; ERM 45%). The authors interpret the lower specificity as a tendency to over-diagnose.

  • Performance drops on external data. On PROVE-AI the overall AUROCs were 0.850 (ADAE), 0.819 (2nd Place), 0.704 (3rd Place) and 0.735 (ERM). At the 95% sensitivity threshold, specificity was 40% (ADAE), 26% (2nd Place), 10% (3rd Place) and 12% (ERM). 3rd Place had the most false positives for every subgroup.

  • ADAE generalizes best of the top three, but the gap to baseline narrows. On PROVE-AI, ADAE differed significantly from 2nd Place in 4 comparisons (Everyone, FST 2 Men, FST 2 Women, FST 1 Age Over 65) and from 3rd Place in 5 (Everyone, FST 2 Men, FST 1 Women, FST 2 Under 65, FST 2 Over 65); ADAE had the superior AUROC in all but 2 of these. Unlike ISIC 2020, several subgroups showed no significant advantage over ERM, and 3rd Place was not significantly better than ERM in any comparison — the authors suggest its high ISIC AUROC likely reflected overfitting to the ISIC data.

  • No significant FST 1 versus FST 2 discrimination gap on PROVE-AI. Stratified by sex, AUROCs ranged from 58% (3rd Place, Women FST 1) to 93% (2nd Place, Men FST 1) for FST 1, and from 64% (3rd Place, Women FST 2) to 85% (ADAE, Men FST 2) for FST 2, all with p-values above 0.05. One marginal result appeared in the FST/sex table (2nd Place, Men, p = 0.09).

  • Every calibration test rejected perfect calibration. Across all experiments the Monte Carlo CUSUM p-value was 0.0, meaning at least one (and possibly several) subgroups had risk systematically over- or underestimated.

  • Age is the dominant calibration feature. Across both datasets and all tested ADAE variants, age ranked second in variable importance in nearly every configuration, behind only the model's own prediction; the authors state ADAE is most likely miscalibrated in terms of age more than any other feature.

  • ADAE is more prone to overestimating than underestimating risk on ISIC 2020. The overestimation test statistic was 127.33 versus 8.45 for underestimation; the authors attribute the gap to the number of samples used (positive versus negative predicted residuals), and note the difference suggests overestimation is the dominant failure mode on that dataset.

  • Calibration varies by model variant and dataset. On ISIC 2020 the individual EfficientNet-NM had underestimation/overestimation statistics of 14.04 and 21.15, while EfficientNet-M had 13.63 and 17.35. On PROVE-AI the same variants produced 10.24/10.53 (NM) and 11.18/11.04 (M). Secondary features differed: body site appears for the PROVE-AI VI rankings, and collection location for ISIC 2020.

  • Including or excluding a risk factor changes calibration. The authors report that whether patient risk factors are included in the audit affects the model's calibration, with some risk factors well calibrated when included and not calibrated when excluded.

Methodology in Plain English

The authors take three already-trained competition models and one simple baseline and do not retrain them on the new data — they run inference with the original ISIC weights, so the models face PROVE-AI as an unseen population.

For discrimination, they compute AUROC overall and for subgroups defined by sex, age, and Fitzpatrick Skin Tone, plus intersections of those subgroups. They compare models to each other using DeLong's test, using the uncorrelated version for "utility" comparisons (different models on the same subgroup) and the correlated version for "group fairness" comparisons (different subgroups under the same model). The significance level is 0.05 and analysis was done in R.

For calibration, they use Feng et al.'s score-based CUSUM test. In plain terms: the audit dataset is split, one part is used to train an ensemble of "residual" models that try to predict the true event rate from metadata and the audited model's own predictions, and the rest is used to compute a test statistic from the products of observed-minus-predicted and predicted-minus-adjusted-predicted residuals. A model is flagged as miscalibrated when the largest of these statistics exceeds a threshold. P-values come from a Monte Carlo simulation of the statistic under perfect calibration. Variable Importance plots then show which features, when perturbed, most change the statistic — pinpointing which subgroups are likely responsible.

The residual ensemble used Kernel Logistic Regression with eight configurations: regularization strengths of 1×10⁻³ (degrees 2, 3, 4, 5) and 1×10⁻² (degrees 2, 3, 4, 5), a maximum of 2000 iterations, L2 regularization, and no weighting of zero-labeled samples. The tolerance δ was set to 0, the CVScore cross-validated variant was used, and the authors added intermediate feature embeddings extracted from the audited model as extra inputs to improve the residual models' estimate of the true event rate.

The ERM baseline is a ResNet-18 trained from scratch with cross-entropy loss and the Adam optimizer, tuned by Bayesian optimization over learning rates in [1×10⁻⁵, 1×10⁻³], weight decays of 1×10⁻⁴ and 1×10⁻⁵, and batch sizes of 256, 512 and 1024, with no augmentation and no metadata. A 10% validation split was held out, the final configuration used a learning rate of 1.475×10⁻⁴, 30 epochs and batch size 256.

Datasets: the ISIC 2020 Challenge dataset comprises a convenience test set of 10,982 public dermoscopy images from six dermatology centers (Barcelona, New York, Vienna, Sydney, Brisbane, Athens), with 10,982 participants, 43% women and 51.6% under age 50; this study computes AUC over the whole dataset rather than the held-out portion used for the original leaderboard. PROVE-AI consists of 603 images collected prospectively at Memorial Sloan Kettering Cancer Center, 95 of them melanoma, with 603 participants (53.7% women, 49.1% under age 65). Because of too few samples for FSTs 4–6, evaluation focuses on FSTs 1 and 2.

Why This Matters

Research impact. The paper argues that AUROC-based group fairness benchmarking, which has dominated recent medical AI audits, cannot detect systematic over- or under-estimation of risk in a subgroup. Adding calibration gives a complementary signal and shows that a model can look fair on discrimination while still producing probabilities that mislead clinicians. It also demonstrates a calibration audit that does not depend on a predefined subgroup list and that handles intersectional membership.

Real-world applications.

  • Clinical decision support for melanoma triage, where a probability score informs whether to biopsy or reassure a patient.
  • Regulatory and hospital procurement review, where a model's behavior on an external population matters as much as its competition leaderboard score.
  • Post-market monitoring of deployed dermatology tools, where a CUSUM-style check could flag drift in calibration over time.
  • Health-equity audits that need to find which demographic combinations drive miscalibration without pre-specifying them.

Industry relevance. The results show that a top leaderboard model can regress toward baseline performance when moved to a new dataset, and that the third-place model's strong ISIC AUROC likely reflected overfitting. For companies deploying medical imaging models, this supports the case that external validation and calibration reporting belong alongside AUROC in evaluation pipelines — and that collecting demographic metadata, including Fitzpatrick skin tone and body site, is a prerequisite for doing that audit at all.

Future Directions

  • Extend the analysis to Fitzpatrick skin tones 4–6. The authors note this was impossible here because those groups were too sparsely sampled, and state the analysis is extendable to datasets with sufficient samples for those groups.
  • Broader external validation. Since all models were evaluated on only two datasets, with PROVE-AI containing 603 images and 95 melanomas, larger and more diverse prospective cohorts would test whether the observed calibration patterns generalize.
  • Investigate the age-driven miscalibration. Age appears at or near the top of the variable importance rankings across nearly every configuration; why the risk estimate loses accuracy with age, and whether retraining or recalibration could fix it, is left open.
  • Determine the clinical cost of miscalibration. The paper establishes that miscalibration exists but does not quantify the downstream harm of a given miscalibration magnitude in a diagnostic workflow.

Target Audience

Researchers and practitioners working on fairness auditing and evaluation of medical imaging models; machine learning engineers deploying or procuring dermatology AI; regulatory and clinical informatics teams who need to interpret model probability outputs rather than just discrimination scores; and fairness methodology researchers interested in subgroup-agnostic calibration tests. Readers without a statistics background will find the AUROC and DeLong test comparisons readable, but the CUSUM test construction requires some patience.

Authors’ abstract

Artificial Intelligence (AI) models have demonstrated expert-level performance in melanoma detection, yet their clinical adoption is hindered by performance disparities across demographic subgroups such as gender, race, and age. Previous efforts to benchmark the performance of AI models have primarily focused on assessing model performance using group fairness metrics that rely on the Area Under the Receiver Operating Characteristic curve (AUROC), which does not provide insights into a model's ability to provide accurate estimates. In line with clinical assessments, this paper addresses this gap by incorporating calibration as a complementary benchmarking metric to AUROC-based fairness metrics. Calibration evaluates the alignment between predicted probabilities and observed event rates, offering deeper insights into subgroup biases. We assess the performance of the leading skin cancer detection algorithm of the ISIC 2020 Challenge on the ISIC 2020 Challenge dataset and the PROVE-AI dataset, and compare it with the second and third place models, focusing on subgroups defined by sex, race (Fitzpatrick Skin Tone), and age. Our findings reveal that while existing models enhance discriminative accuracy, they often over-diagnose risk and exhibit calibration issues when applied to new datasets. This study underscores the necessity for comprehensive model auditing strategies and extensive metadata collection to achieve equitable AI-driven healthcare solutions. All code is publicly available at https://github.com/bdominique/testing_strong_calibration.

Read the original paper