Skip to content
AI.info

Research

Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks

Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks Overview Research area: Uncertainty quantification and trustworthy machine learning — specifically split conform

Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks
arXiv
2511.18562
Published
2025-11-23
Authors
Xunlei Qian, Yue Xing

AI summary

Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks

Overview

Research area: Uncertainty quantification and trustworthy machine learning — specifically split conformal prediction (CP), adversarial robustness, and distribution shift.

Technical level: Advanced. The paper is theory-first, with formal coverage bounds, a tolerance-band theorem, and a prediction-set-size bound, followed by image-classification experiments on adversarially attacked data.

Scope (one sentence): The paper analyzes how adversarial perturbations applied during the calibration step of split conformal prediction change coverage validity and prediction-set size when test inputs are also adversarially perturbed, and how adversarial training during model fitting affects the resulting set sizes.

What This Paper Is About

Conformal prediction gives prediction sets with finite-sample coverage guarantees, but those guarantees rely on exchangeability between calibration and test data — an assumption that adversarial perturbations deliberately break. This paper asks three linked questions: how does the strength of an attack applied to the calibration set change coverage when the test set is attacked; can a nonzero calibration-time attack be chosen so that target coverage is held within a chosen tolerance over a range of test-time attack strengths; and does adversarial training during model fitting shrink the conformal prediction sets.

Key Contributions

  1. A theoretical analysis (Theorem 1) showing that the prediction-coverage gap relative to the target level depends on the calibration attack strength, and that coverage is a monotone non-decreasing function of calibration attack strength when the test-time attack range is held fixed.
  2. A result (Theorem 2) showing that an appropriately chosen calibration attack lets split CP keep coverage inside a tolerance band around the target — for example 87%–93% around a 90% target — across a contiguous range of test-time attack strengths, with an explicit expression for the length of that interval.
  3. A bound (Theorem 3) showing that adversarial training reduces the expected split-conformal prediction-set size relative to clean training, under a uniform-residual assumption on the model's predictive distribution.
  4. An empirical study across CIFAR-10, CIFAR-100, TinyImageNet, and MNIST with ResNet-50d and Vision Transformer (ViT) backbones, confirming the monotonicity, the tolerance-band behavior, and the set-size reduction.

Main Findings

  • Coverage moves monotonically with calibration attack strength. In Theorem 1, for a fixed test-time attack strength, the coverage of the split CP set produced from an attacked calibration set satisfies a bound of the form |P(y_test ∈ C_{ε_cal}(x_test + ε_test A_test)) − (1−α)| ≤ e_train + d_cal + e_cal + ((2ε_test − ε_cal)||∇f_{x_test}|| − ε_cal·c) g(Q_{1−α}(f_{y_cal}(x_cal))) + o(ε_cal² + ε_test²), where g is the density of the true-label probability. Larger ε_cal induces larger prediction sets and hence a smaller accuracy shortfall.

  • The calibration attack shifts, but does not stretch, the robust test-time interval. Theorem 2 states that P(y_test ∈ C(x_test + ε_test A_test)) lies in [1−α−β, 1−α+β] for a range of ε_test, with interval length 2β / (2||∇f_{x_test}|| g(Q_{1−α}(f_{y_cal}(x_cal)))). Increasing ε_cal moves that interval toward larger attack levels and decreasing it moves the interval toward smaller ones; the length is governed by the tolerance β and the model's local sensitivity.

  • Adversarial training yields smaller prediction sets. Theorem 3 bounds the expected set size as E(|C(x_test)|) ≤ 1 − α ± h(ε_cal, ε_test) + (K−1)(1 − (1−Q)/(1−P_{y,true}))^{K−2}, where Q is the (1−α)-quantile of the HPS score and P_{y,true} is the true-label probability. Because the bound increases in (1−Q)/(1−P_{y,true}), an adversarially trained model with a lower ratio yields a smaller expected set size than a cleanly trained one.

  • Experiments match the theory. With adversarial training at ε_train = 4/255, test accuracy was nonincreasing in ε_test for ε_cal ∈ {8/255, 16/255}, and for fixed ε_test a larger ε_cal gave higher accuracy (Theorem 1). With ε_cal fixed, the ε_test interval keeping performance inside the 88%–92% band widened as ε_cal increased, while prediction-set sizes stayed nearly unchanged (Theorem 2). Adversarially trained models produced substantially smaller prediction sets than cleanly trained models across CIFAR-10/100, MNIST, and TinyImageNet (Theorem 3).

  • RSCP did not hold coverage in a preliminary check. The authors note that preliminary trials on CIFAR-100 showed Randomized Smoothed Conformal Prediction (RSCP), under an ℒ∞ attack with strength 8/255 on the test set, failed to maintain the desired coverage, which is why they focus on their own approach.

  • Setup detail. The theory is derived under ℓ2-bounded perturbations for algebraic clarity, and the proofs are carried out under that assumption; the experiments use ℓ∞ attacks (ℓ∞-FGSM) for adversarial training, calibration-time attacks, and test-time attacks. The authors state that at comparable budgets ℓ2 perturbations left clean models with high accuracy and were not discriminative, which is why ℓ2 results are not reported.

Methodology in Plain English

The pipeline has three stages. First, a classifier — ResNet-50d or ViT — is trained with adversarial training using attack strength ε_train. Second, the calibration set is attacked with a deliberately chosen strength ε_cal, and split conformal prediction builds the prediction set using the HPS nonconformity score S(x, y) = 1 − f_y(x), with the threshold Q set to the (1−α)(1 + 1/|I₂|)-quantile of the calibration nonconformity scores. Third, test samples may be attacked at an unknown strength ε_test. Data is split 60% training, 20% calibration, and 20% test.

The theoretical analysis builds on a two-sided error bound from prior work (|P(Y_{n+1} ∈ C(X_{n+1})) − 1 − α| ≤ e_train + d_cal + e_cal under exchangeability), then adds a first-order term describing how the calibration and test perturbations move the probability that the true label's score clears the threshold. The density g of the true-label probability and its second-order differentiability are assumed so this term can be Taylor-expanded. For the set-size result, the authors use a supporting example from prior work showing adversarially trained two-layer networks (width m = polylog(d), T = Θ(poly(d)/η) iterations) reach a robust global minimum where cleanly trained ones do not, then split the prediction set into a true-label component and a uniform-residual component over the remaining K−1 classes.

Empirically, the authors sweep a grid: for a calibration level of 8/255, training ε ranges over {0, 4, 8, 12, 16}/255 and test ε over {2–14}/255; for a calibration level of 16/255, the same training range is used with test ε over {10–22}/255. All models were trained for 40 iterations on a single NVIDIA H200 GPU, with a 90% target coverage and tolerance β = 2%. The paper mentions that across its parameter grid it explores prediction accuracy "across five different datasets," while four datasets (CIFAR-10, CIFAR-100, TinyImageNet, MNIST) are named in the setup. Additional results and the ResNet-50d counterparts are placed in the paper's appendix.

Why This Matters

Impact on research. The work gives a quantitative rather than purely empirical account of what happens to conformal validity when the calibration set itself is perturbed — a setting that standard exchangeability-based analyses do not cover. It converts the calibration attack strength into a tuning knob with a predictable effect on coverage, and it gives a separate account of why adversarial training improves conformal efficiency, linking adversarial robustness research to the coverage-versus-set-size tradeoff in conformal prediction.

Real-world applications:

  • Clinical decision support, where physicians deploy complex diagnostic models and need calibrated uncertainty (the paper cites survival analysis and chest X-ray report generation as motivating high-stakes examples).
  • Any deployment where inputs may be deliberately corrupted, such as content-moderation or fraud detection pipelines facing evasion attempts.
  • Autonomous or safety-critical perception systems where inputs can be adversarially perturbed and a coverage guarantee must still be trusted.
  • Scientific and engineering pipelines that use prediction sets for downstream decisions under measurement corruption or distribution shift.

Industry relevance. Practitioners deploying conformal prediction in adversarial settings get a concrete prescription — pick a nonzero calibration attack to buy coverage stability across a band of test-time attack strengths — and a reason to combine adversarial training with conformal calibration instead of using either alone.

Future Directions

  • Extending the analysis beyond the ℓ2-bounded perturbation assumption used in the proofs to match the ℓ∞ threat model used in experiments.
  • Removing or weakening the uniform-residual assumption in Theorem 3, which assumes the model spreads probability mass evenly across incorrect classes and is described as standard but conservative.
  • Understanding the gap between anticipated and actual test-time attack strength: the theory shows calibration shifts the tolerance interval but does not change its length, so how to choose ε_cal when the test attack distribution is unknown or adaptive remains open.
  • Comparing against adversarially robust conformal baselines more systematically — the paper reports only a preliminary CIFAR-100 check in which RSCP failed to hold coverage under an ℓ∞ attack with strength 8/255.
  • Reconciling the theoretical ℓ2 analysis with empirical ℓ∞ results across a wider range of architectures and datasets, since the main experiments cover four named datasets and two architectures with additional results deferred to the appendix.

Target Audience

Researchers and graduate students working on conformal prediction, distribution shift, or adversarial robustness; practitioners who deploy prediction sets in safety-critical or adversarial environments; and methodologists interested in how calibration-time design choices propagate to finite-sample coverage guarantees. A reader needs comfort with conformal prediction notation, nonconformity scores, and first-order perturbation analysis to follow the theorems, though the empirical sections are accessible on their own.

Authors’ abstract

Conformal prediction (CP) provides distribution-free, finite-sample coverage guarantees but critically relies on exchangeability, a condition often violated under distribution shift. We study the robustness of split conformal prediction under adversarial perturbations at test time, focusing on both coverage validity and the resulting prediction set size. Our theoretical analysis characterizes how the strength of adversarial perturbations during calibration affects coverage guarantees under adversarial test conditions. We further examine the impact of adversarial training at the model-training stage. Extensive experiments support our theory: (i) Prediction coverage varies monotonically with the calibration-time attack strength, enabling the use of nonzero calibration-time attack to predictably control coverage under adversarial tests; (ii) target coverage can hold over a range of test-time attacks: with a suitable calibration attack, coverage stays within any chosen tolerance band across a contiguous set of perturbation levels; and (iii) adversarial training at the training stage produces tighter prediction sets that retain high informativeness.

Read the original paper