Skip to content
AI.info

Research

Torch-Uncertainty: A Deep Learning Framework for Uncertainty Quantification

Overview Research area: Uncertainty quantification (UQ) for deep learning; scientific software / benchmarking frameworks. Technical level: Intermediate — the paper is a library-and-benchmark paper rat

Torch-Uncertainty: A Deep Learning Framework for Uncertainty Quantification
arXiv
2511.10282
Published
2025-11-13
Authors
Adrien Lafage, Olivier Laurent, Firas Gabetni, Gianni Franchi

AI summary

Overview

Research area: Uncertainty quantification (UQ) for deep learning; scientific software / benchmarking frameworks. Technical level: Intermediate — the paper is a library-and-benchmark paper rather than a new algorithm, but it assumes familiarity with calibration, out-of-distribution detection, and ensemble methods. Scope: The paper introduces Torch-Uncertainty, an open-source PyTorch and Lightning library for training and evaluating deep neural networks with uncertainty quantification, and reports classification and semantic-segmentation benchmarks produced with it.

What This Paper Is About

Deep neural networks produce predictions without reliable confidence estimates, which makes them risky to deploy in high-stakes settings. Many UQ techniques exist (ensembles, Bayesian neural networks, post-hoc calibration, conformal prediction, and others), but there has been no single, unified tool for training models with these methods and evaluating them consistently across tasks. The authors build Torch-Uncertainty to fill that gap: one library covering many UQ methods, many metrics, many datasets, and multiple learning tasks.

Key Contributions

  1. Torch-Uncertainty itself — described by the authors as the first unified, extensible, domain-general and evaluation-centric PyTorch-based library for uncertainty quantification in deep learning. It is built on PyTorch and Lightning, released under the Apache-2.0 license, and hosted at https://github.com/ENSTA-U2IS-AI/Torch-Uncertainty.
  2. A modular implementation of UQ methods across six of the seven method families the authors identify (ensemble-based, Bayesian, post-hoc, data-augmentation-based, deterministic, and interval/conformal), usable across classification, regression, segmentation, and pixel-wise regression.
  3. Reproducible benchmarks of these methods on standard datasets and tasks, including a ViT-B/16 classification benchmark on ImageNet and a UNet segmentation benchmark on MUAD.
  4. Pretrained models and educational material — a benchmarked model zoo on Hugging Face, datasets on Zenodo, interactive tutorials, and documentation.

Main Findings

  • Library scope: As of release v0.7.0, the library reaches nearly complete unit test coverage (around 98%), uses ruff for code-style compliance and pytest for unit tests, and counts 26 distinct metrics spanning seven task categories (classification, out-of-distribution detection, selective classification, calibration, diversity, regression/depth prediction, and segmentation), plus efficiency metrics (number of parameters and floating point operations). For comparison, the paper states Lightning-UQ-Box provides nine metrics in four categories, and that all other surveyed libraries expose three metrics or fewer.
  • Method coverage: In Table 1, Torch-Uncertainty is marked as implementing Deep Evidential, Beta-Gaussian Regression, Deep Ensembles, BatchEnsemble, Masksembles, MIMO, Packed-Ensembles, Snapshot Ensemble, Variational BNN, LP-BNN, SWA, SWAG, SGLD, SGHMC, conformal classification, temperature scaling, test-time augmentation, Laplace approximation, MC-Dropout, MCBatchNorm, and 15 different OOD evaluation methods. The authors deliberately do not implement Gaussian-process families, citing difficulty scaling them to larger models.
  • Classification benchmark setup: A ViT-B/16 was pre-trained on ImageNet-21k and fine-tuned on ImageNet-1k; the whole pipeline was repeated three times with different random seeds to form a deep ensemble. Benchmarked variants were a single model, single model plus temperature scaling, Packed Ensemble, MiMo, and Deep Ensemble (with and without temperature scaling).
  • Deep Ensembles win on accuracy and likelihood: The Deep Ensemble reaches the best overall accuracy (82.19% versus 80.67% for the single model) and the lowest Brier/NLL (0.25/0.65). After temperature scaling its calibration improves to ECE = 0.01%, with AURC = 4.41%, Cov@5Risk = 67.9%, and Risk@80Cov = 8.49%.
  • Compact ensembles trade quality for efficiency: Packed Ensemble and MiMo reach comparable accuracies (79.2–80.59%) but weaker Risk@80Cov (10.9% and 9.6% respectively).
  • Deep Ensembles also lead on out-of-distribution detection: Highest FarOOD AUROC (92.05%) and lowest FPR95 (33.05%), beating the single model by +1.3 AUROC and −4.9 FPR95. On NearOOD it reaches 78.8% AUROC and 61.7% FPR95 (78.8% / 61.7%); MiMo remains competitive at 78.1% / 62.8%.
  • Segmentation benchmark setup: UNet backbones evaluated on MUAD (3420 training samples, 492 validation, 551 in-distribution test, 1668 out-of-distribution test; 15 in-distribution classes and 6 out-of-distribution classes). All ensembles used 4 subnetworks; MC Dropout used 10 forward passes; Packed-Ensembles used α = 2 and γ = 1; MIMO used ρ = 0.5; Masksemble used scale 2.0. Results are averaged over three runs.
  • Deep Ensembles best on segmentation, but not best on calibration: Deep Ensembles achieves the top mIoU (74.93%), mAcc (88.86%), and pixAcc (94.29%) on MUAD, ahead of the baseline (71.55 / 87.65 / 93.59) and Packed-Ensembles (71.87 / 86.77 / 93.65). However, it is less well calibrated than the baseline (Brier 0.09 vs 0.10 but ECE 1.58% vs 0.51%, aECE 2.14% vs 0.42%); the authors attribute this to training-time data augmentation and use it to argue for automated calibration assessment.
  • Selective classification and OOD detection on MUAD: Deep Ensembles again leads on AURC (0.63), AUGRC (0.56), Cov@5Risk (98.53%), Risk@80Cov (0.93%), AUPR (22.45%), and AUROC (84.03%). Packed-Ensembles performs well using only 25% of Deep Ensembles' parameters, which the authors read as evidence that the subnetworks may be under-parameterised; they suggest a higher α (e.g., α = 3) would help, noting that α = 4 would give Packed-Ensembles the same parameter count as Deep Ensembles at ensemble size 4. BatchEnsemble has the best FPR95 (48.00%) but the weakest segmentation capability (mIoU 64.88%).
  • Dataset breadth: The library ships with 12 corrupted vision dataset variants (from MNIST-C and CIFAR10/100-C,H,N up to ImageNet-A/C/O/R and TinyImageNet-C), OOD vision sets (the paper says "six popular out-of-distribution sets" but names Places365, Textures, SVHN, iNaturalist, NINCO, SSB-hard, and OpenImages-O — seven names), dense-prediction sets (CamVid, Cityscapes, MUAD for segmentation; Fractals, Frost, KITTI-Depth, NYUv2 for depth/texture or synthetic images), and UCI tabular loaders (BankMarketing, Dota2, HTRU2, OnlineShoppers, SpamBase, plus a UCI-Regression loader that cycles through 9 regression datasets). The paper summarises this as 27 plug-and-play datasets. OOD-vision splits follow those defined in the OpenOOD library.
  • Training utility: A CompoundCheckpoint object saves the best checkpoint according to a combination of validation metrics, addressing the fact that the best epoch differs per metric (illustrated with a UNet on MUAD).
  • Reported but not detailed in the provided content: The paper states that Appendix E and Appendix I report benchmarks on regression and time-series classification respectively; those results are not included in the content above.

Methodology in Plain English

The authors take an engineering-plus-empirical approach. First, they survey existing UQ toolkits (TorchCP, Fortuna, TorchUQ, MAPIE, BLiTZ, Bayesian-Torch, Lightning-UQ-Box, Uncertainty Toolbox, GPyTorch, Uncertainty Baselines, NeuralUQ) and position their library against them on three axes: domain generality, modular composition of UQ techniques, and evaluation-centric design. Second, they build the library around task-specific "routines" (classification, regression, pixel-wise regression, segmentation) that wrap PyTorch models into Lightning training loops, computing uncertainty-aware metrics during validation and saving the best checkpoints. Third, they demonstrate the framework by running two full benchmarks: a two-stage ViT-B/16 pipeline (ImageNet-21k pre-training, then ImageNet-1k fine-tuning) run three times with different seeds, and a UNet segmentation study on MUAD with several ensemble variants, all averaged over three runs. Every method is scored on the same metric battery — accuracy, Brier, NLL, ECE, aECE, AURC, AUGRC, Cov@5Risk, Risk@80Cov, AUPR, AUROC, FPR95, mIoU, mAcc, pixAcc — so methods can be compared directly.

Why This Matters

Impact on research: UQ research is fragmented across incompatible codebases, which makes comparing methods hard and reproducing results harder. By providing one modular routine-based architecture, 26 metrics, 27 datasets, pretrained checkpoints, and a test suite with around 98% coverage, Torch-Uncertainty lowers the cost of running comparable experiments and makes it easier to prototype hybrids (for example, an ensemble of Laplace approximations, or an ensemble of MC Dropout models) by swapping layers and post-processing steps.

Real-world applications:

  • Medical imaging, where overconfident errors can lead to inappropriate treatment decisions.
  • Autonomous driving, where uncertain perception should trigger caution or handover.
  • Finance, where predictions need reliable confidence before being acted on.
  • General safety-critical deployment, where uncertainty scores drive human intervention, deferred decisions, or risk flagging.

Industry relevance: The library targets practitioners as well as researchers — the Hugging Face model zoo enables plug-and-play testing, the Zenodo datasets remove data-plumbing work, and the routines fit into the widely used PyTorch/Lightning stack. The benchmark finding that ensembles can be well-calibrated on one deployment-relevant metric while missing calibration on another (as with Deep Ensembles on MUAD) is directly useful to teams deciding how much evaluation is enough.

Future Directions

  • Close the coverage gap: The library currently supports six of the seven UQ families, and the authors note maintenance and development costs limit comprehensiveness. Expanding method coverage without breaking modularity is an open engineering problem.
  • Broaden beyond the four tasks and computer vision: The authors state the library is specialised on computer vision and limited to four main tasks (classification, regression, segmentation, pixel-wise regression); extending to other modalities and tasks is a stated goal.
  • Resolve the calibration/accuracy trade-off observed in the segmentation benchmark: The finding that Deep Ensembles is less calibrated than a plain baseline on MUAD — attributed to training-time augmentation — raises the question of how to obtain both strong segmentation and strong calibration.
  • Tune compact ensembles: The segmentation results suggest Packed-Ensembles may be under-parameterised at α = 2 with ensemble size 4, leaving open whether larger α values recover Deep Ensemble performance at lower cost.
  • Reduce bug risk from modularity: The authors acknowledge the modular structure is more prone to compatibility bugs between components and rely on unit and integration tests to mitigate this.

Target Audience

Researchers and graduate students working on uncertainty quantification, calibration, selective

Authors’ abstract

Deep Neural Networks (DNNs) have demonstrated remarkable performance across various domains, including computer vision and natural language processing. However, they often struggle to accurately quantify the uncertainty of their predictions, limiting their broader adoption in critical real-world applications. Uncertainty Quantification (UQ) for Deep Learning seeks to address this challenge by providing methods to improve the reliability of uncertainty estimates. Although numerous techniques have been proposed, a unified tool offering a seamless workflow to evaluate and integrate these methods remains lacking. To bridge this gap, we introduce Torch-Uncertainty, a PyTorch and Lightning-based framework designed to streamline DNN training and evaluation with UQ techniques and metrics. In this paper, we outline the foundational principles of our library and present comprehensive experimental results that benchmark a diverse set of UQ methods across classification, segmentation, and regression tasks. Our library is available at https://github.com/ENSTA-U2IS-AI/Torch-Uncertainty

Read the original paper