Research
Conformal Prediction Sets for Instance Segmentation
Conformal Prediction Sets for Instance Segmentation Overview Research area: Uncertainty quantification for computer vision, specifically conformal prediction applied to instance segmentation. Technica

- arXiv
- 2602.10045
- Published
- 2026-02-10
- Authors
- Kerri Lu, Dan M. Kluger, Stephen Bates, Sherrie Wang
AI summary
Conformal Prediction Sets for Instance SegmentationOverview
Research area: Uncertainty quantification for computer vision, specifically conformal prediction applied to instance segmentation.
Technical level: Advanced. The paper combines conformal prediction theory (asymptotic and finite-sample coverage guarantees, set cover and dominating set problems) with applied instance segmentation experiments.
Scope: The paper introduces and empirically validates a conformal algorithm that returns a small set of candidate instance masks for a pixel query, with a provable guarantee that at least one mask in the set has high Intersection-Over-Union (IoU) with the true object, capturing "structural" uncertainty such as whether adjacent regions should be one object or several.
What This Paper Is About
Instance segmentation models output masks that look good on average, but their confidence scores are not calibrated and there is no guarantee that any predicted mask is close to the ground-truth object. Existing conformal methods for segmentation only perturb a single predicted mask (for example by dilating it or expanding its boundary), which cannot represent structural ambiguity such as a model merging two fields into one or splitting one object into two. This paper builds confidence sets containing several qualitatively different masks for a given image and pixel-coordinate query, such that with probability at least 1 − α, at least one mask in the set has IoU above a user-specified threshold τ with the true instance.
Key Contributions
-
Identification of structural uncertainty not captured by existing conformal approaches that rely on boundary perturbations (dilation or boundary expansion) of a single predicted mask.
-
A conformal prediction formulation for set-valued outputs: rather than selecting one model configuration, the method selects a set of configurations whose predictions jointly achieve coverage, which remains feasible even when no single parameter value reaches the desired IoU threshold.
-
A practical algorithm instantiation using a sweep over a tunable model parameter plus set-cover-based selection (greedy set cover followed by brute-force refinement), with an optional duplicate-removal step that produces prediction sets whose sizes adapt to query difficulty.
-
Asymptotic and finite-sample versions of the guarantee, plus evaluations on three real segmentation domains: agricultural field delineation, cell segmentation, and vehicle detection.
The paper states it is the first to capture structural uncertainty in instance segmentation by constructing confidence sets of diverse segmentation predictions.
Main Findings
-
Target coverage is attained, and beats baselines. For field delineation (τ = 0.7), conformal coverage is 0.797 against a target of 0.8, versus 0.663 for the naive best parameter baseline. For cell segmentation (τ = 0.75), conformal coverage is 0.835 against a target of 0.8, versus 0.803 for the baseline. For vehicle detection (τ = 0.8), conformal coverage is 0.843 against a target of 0.9, versus 0.661 for the baseline.
-
Reported coverage gains. The paper describes coverage increasing from 66.3% to 79.7% for field delineation, from 80.3% to 83.5% for cell segmentation, and from 66.1% to 84.3% for vehicle detection.
-
Re-calibrated IoU thresholds stay near or above the original target. After duplicate removal, θ̃ is 0.696 for fields (τ = 0.7), 0.760 for cells (τ = 0.75), and 0.842 for vehicles (τ = 0.8). The paper notes that in all three examples θ̃ > τ or θ̃ ≈ τ, so duplicate removal does not significantly affect the original guarantee.
-
Prediction sets adapt to query difficulty. After duplicate removal, most sets contain three or fewer masks. Field delineation produces the largest sets, going up to 5 potential masks, reflecting under- and over-segmentation hypotheses. For cells, most sets collapse to size one or two, with a few sets of size zero because the model classifies the queried pixel as a non-cell.
-
Domain difficulty differs. Field delineation has the weakest guarantee and is described as the most ambiguous of the three tasks, while SAM vehicle segmentation has the strongest. Cell segmentation improves least, which the paper attributes to most errors in that domain being boundary-local.
-
Feasibility ceilings exist. For field delineation it was not possible to construct a 90% confidence set for IoU > 0.8, because even taking the best T per query, more than 10% of predictions had IoU < 0.8. The chosen operating points were (α = 0.2, τ = 0.7) for fields, (α = 0.2, τ = 0.75) for cells, and (α = 0.1, τ = 0.8) for vehicles.
-
The dilation baseline underperforms. The morphological dilation method of Mossina and Friedrich (2025) only guarantees that the true mask is contained within the dilated mask with high probability, and the paper reports that it results in low IoU coverage, undersegmentation, and lack of flexibility compared with their method.
-
Best single-parameter values on the calibration sets were T = 0.24 for field delineation, T = 0 for cell segmentation, and T = (1, 0.05) for vehicle detection. Because no single value covers all queries, the naive best parameter baseline cannot match the conformal sets.
Methodology in Plain English
The method assumes access to a model that produces masks and exposes a tunable parameter T — for example, the threshold in watershed segmentation, a cell-extent threshold, or a mask index plus probability threshold in Segment Anything. The researchers sweep a grid of k values of T over a calibration set of image-query pairs with known ground-truth masks. For each calibration point and each parameter value, they compute the IoU between the predicted mask and the truth, and record which calibration points exceed the target IoU τ for each parameter value. The core step is then choosing the smallest set of parameter values whose successful calibration points collectively cover at least (1 − α)·n of the calibration set. This is a variant of the NP-hard Set Cover problem, so they solve it with the polynomial-time greedy set cover algorithm, optionally refining by brute-forcing over smaller subsets to find a truly minimal cover. At test time, they run the model with each selected parameter value and return those masks as the confidence set.
To avoid redundant outputs, they post-process the set by removing near-duplicate masks — masks within IoU η of one another are considered duplicates. Because this post-processing breaks the original calibration, they re-calibrate by computing, for every calibration point, the best IoU inside its own duplicate-removed set, and taking the α-quantile of those scores as the new threshold θ̃. This yields an asymptotic guarantee (Theorem 2.1) that with probability at least 1 − α, some mask in the test set exceeds θ̃ in IoU with the truth. A separate finite-sample variant, given in the appendix, randomly splits the calibration data, uses a greedy procedure on the first half to rank parameters, and applies Conformal Risk Control on the second half. The experiments in the main text use the asymptotic version, with η = 0.9 for duplicate removal throughout.
The authors note a practical limitation: not all (α, τ) pairs are achievable. If more than α·n calibration points have no parameter value achieving IoU above τ, the algorithm's line 8 fails, and the user must increase α, lower τ, or broaden the parameter grid. They also note that choosing α and τ by inspecting calibration histograms is a mild form of double-dipping in the calibration data.
Why This Matters
Impact on research. The paper argues that prior conformal segmentation work — Conformal Risk Control (Angelopoulos et al., 2022), Learn Then Test (Angelopoulos et al., 2025b), and Risk-Controlling Prediction Sets (Bates et al., 2021) — selects a single model configuration and can fail when no single parameter value achieves the target. It also notes IoU loss is neither monotone nor near-monotone, which blocks the use of CRC and RCPS directly, whereas this method does not require monotonicity. By reframing the object of conformal calibration as a set of configurations rather than one, the paper opens a route to coverage guarantees in settings where a single model output is structurally incapable of being correct.
Real-world applications (as presented in the paper):
- Agricultural field delineation, where farmers select their fields in satellite imagery and the model must decide whether adjacent regions are one field or several.
- Cell segmentation, where boundary-local errors dominate and instance delineation is critical for biological analysis.
- Vehicle detection on Cityscapes imagery, where SAM's top mask can miss or mis-delineate cars and alternative hypotheses in the set recover them.
- Interactive segmentation with point prompts, as in Segment Anything, which returns three masks per query without any guarantee that any of them is correct or that the predicted IoUs are accurate.
Industry relevance. The approach turns an existing segmentation model's tunable parameter into a calibration knob, so practitioners do not need to retrain a model to obtain statistical guarantees. The paper explicitly frames this as quantifying how far current models are from reliability levels needed in practice, noting that conformal prediction can certify but not overcome the limits of the underlying model. The adaptive set sizes — usually three or fewer masks, expanding only for ambiguous queries — matter for interfaces where users must review outputs.
Future Directions
-
Scaling duplicate removal and set selection. Both the set cover step and the dominating-set duplicate removal are NP-hard; the paper uses brute-force refinement that it says is tractable only because the resulting sets are small, and notes that larger sets would require approximation or heuristic algorithms.
-
Achieving finite-sample rather than asymptotic guarantees with minimal sets. The finite-sample version in the appendix splits the calibration data to preserve exchangeability and restricts the range of possible prediction sets, so it cannot guarantee minimal-size outputs. Closing the gap between the finite-sample and asymptotic versions is an open direction.
-
Removing double-dipping in the choice of α and τ. The authors acknowledge that selecting α and τ by inspecting calibration histograms is a form of double-dipping, and note that prespecifying them can cause the algorithm to return an error when the pair is infeasible.
-
Broadening the pool of models and parameters. When a feasible (α, τ) pair does not exist, the paper suggests considering a superset of segmentation model parameters or a broader collection of models, which points to extending the method beyond a single model's parameter sweep.
Sensitivity analyses varying α, τ, η, the parameter-space size k, and the calibration set size n are reported in Appendix K.
Target Audience
Researchers and practitioners working on uncertainty quantification, conformal prediction, or reliable computer vision — particularly those applying segmentation models in domains such as remote sensing, biology, and autonomous driving where a wrong mask has real consequences. It is also relevant to readers interested in the theory of conformal methods, since it presents both asymptotic and finite-sample guarantees and discusses how its construction departs from standard conformal prediction by not using a non-conformity score or exchangeability arguments. A working familiarity with segmentation metrics such as IoU and with conformal coverage guarantees will help, though the paper defines its key quantities explicitly.
Authors’ abstract
Current instance segmentation models achieve high performance on average predictions, but lack principled uncertainty quantification: their outputs are not calibrated, and there is no guarantee that a predicted mask is close to the ground truth. To address this limitation, we introduce a conformal prediction algorithm to generate adaptive confidence sets for instance segmentation. Given an image and a pixel coordinate query, our algorithm generates a confidence set of instance predictions for that pixel, with a provable guarantee for the probability that at least one of the predictions has high Intersection-Over-Union (IoU) with the true object instance mask. We apply our algorithm to instance segmentation examples in agricultural field delineation, cell segmentation, and vehicle detection. Empirically, we find that our prediction sets vary in size based on query difficulty and attain the target coverage, outperforming baselines (naive best parameter and morphological dilation-based methods). We provide versions of the algorithm with asymptotic and finite sample guarantees. Our work is the first to capture structural uncertainty in instance segmentation by constructing confidence sets of diverse segmentation predictions.