Research
Beyond Binary Out-of-Distribution Detection: Characterizing Distributional Shifts with Multi-Statistic Diffusion Trajectories
Overview Research area: Out-of-distribution (OOD) detection and distribution-shift characterization in machine learning, with a focus on denoising diffusion probabilistic models (DDPMs). Technical lev

- arXiv
- 2510.17381
- Published
- 2025-10-20
- Authors
- Achref Jaziri, Martin Rogmann, Martin Mundt, Visvanathan Ramesh
AI summary
Overview
Research area: Out-of-distribution (OOD) detection and distribution-shift characterization in machine learning, with a focus on denoising diffusion probabilistic models (DDPMs).
Technical level: Intermediate to Advanced. The paper combines an empirical benchmarking study, a theoretical non-identifiability result, and a novel generative-model-based scoring pipeline.
Scope: The paper argues that scalar OOD scores cannot distinguish between types of distributional shift, proves this limitation, and proposes DISC (Diffusion-based Statistical Characterization), a multi-statistic embedding built from diffusion denoising trajectories, evaluated on image and tabular benchmarks.
What This Paper Is About
Modern OOD detectors compress all evidence about an input into a single scalar outlier score, so a corrupted image, an adversarially perturbed image, and an image of a completely unseen class all end up looking the same: "not in-distribution." The paper's goal is to move beyond this binary decision by extracting a multi-dimensional feature vector from the iterative denoising trajectory of a diffusion model, which captures statistical discrepancies across multiple noise levels and multiple complementary metrics, and then uses that vector both for standard OOD detection and for classifying which OOD family a sample belongs to.
Key Contributions
-
An empirical demonstration that scalar detectors conflate shift types. The authors show that Mahalanobis (with and without ODIN), Energy-based, and SHE scores produce OOD histograms that overlap heavily with one another on a shared CIFAR-10 in-distribution base, even while standard ID-vs-OOD AUROC stays high.
-
A theoretical non-identifiability result. Extending prior work (Zhang et al., 2021), the authors prove that if the conditional law P(X | φ_p(X) = φ) is non-degenerate on a set of nonzero measure, then there exist two distinct distributions Q₁ and Q₂, both different from P and from each other, that induce the same marginal distribution over the test statistic φ_p(X). Any test depending only on φ_p(x) therefore has power equal to its false positive rate.
-
The DISC framework. DISC replaces a single anomaly score with the embedding s(x) ∈ ℝ^{NM}, built from M complementary metrics evaluated at N noise levels of a diffusion trajectory.
-
A new multi-OOD evaluation protocol. Alongside standard Avg AUROC, the paper evaluates OOD-type classification using unsupervised K-means (Clust. Acc) and a supervised two-layer MLP (Sup. Acc), and includes an "All detectors" baseline that concatenates all scalar baseline scores into one vector to control for the dimensionality advantage of DISC.
Main Findings
-
OOD scores overlap across shift types. For models trained on CIFAR-10, different OOD families each shift away from the ID mode, but their histograms overlap heavily with one another. The relative ordering is consistent across detectors: CIFAR-FGSM and CIFAR-Flip cluster closer to the ID peak, while Butterflies and MNIST are scored as more novel.
-
Binary detection works; type classification does not. Standard ID-vs-OOD AUROC remains high across detectors, but accuracy for identifying which OOD family a sample belongs to drops markedly under both K-means and MLP protocols.
-
DISC on CIFAR-10. Avg AUROC 0.8261 ± 0.0108, Clustering Accuracy 0.5120 ± 0.0110, Supervised Accuracy 0.7131 ± 0.0180. For comparison, the strongest baseline AUROC reported is Mahalanobis at 0.8415 ± 0.011, with Maha+ODIN at 0.8381 ± 0.0111. The "All detectors" concatenation achieves 0.4907 ± 0.0210 clustering accuracy and 0.7084 ± 0.0184 supervised accuracy (AUROC is reported as "–" for that row).
-
DISC on ImageNet. Avg AUROC 0.7404 ± 0.0119, Clustering Accuracy 0.4776 ± 0.0194, Supervised Accuracy 0.6203 ± 0.0163. The DDPM baseline from Graham et al. (2023) reaches 0.6502 ± 0.0121 AUROC, 0.3275 ± 0.0165 clustering accuracy, and 0.3905 ± 0.0172 supervised accuracy, while "All detectors" reaches 0.2343 ± 0.0130 and 0.3374 ± 0.0152 on the two multi-OOD metrics.
-
MSE-only diffusion already helps. On ImageNet, a DDPM using only MSE reconstruction error matches or exceeds several baselines in Avg AUROC and provides clear gains in multi-OOD discrimination; the full DISC embedding improves further.
-
Tabular results. On a subset of datasets from the ADBench benchmark (Han et al., 2022), DISC achieves an average AUROC of 0.86, surpassing the strongest comparable reconstruction-based baseline (DDPM with MSE-iForest, 0.84).
-
All results are reported as mean ± std over 5 seeds, with per-dataset scores deferred to the Appendix.
-
DISC changes the winning detector. Notably, on ImageNet the baseline detectors score near chance on the multi-OOD metrics (e.g., Mahalanobis at 0.1869 ± 0.0112 clustering accuracy), while DISC reaches 0.4776.
Methodology in Plain English
The authors start from a score-based generative model trained with the Variance Preserving formulation used in DDPMs. Such a model learns a family of denoisers D_σ, one for each noise level σ. Give the model a clean image x₀, corrupt it with Gaussian noise at level σ to get x_σ, and ask the denoiser to reconstruct it; the difference between x_σ and D_σ(x_σ) tells you something about how well that image fits the learned distribution at that level of abstraction.
DISC's core idea is to take that single comparison and turn it into many comparisons. At each noise level, four metrics are computed: Mean Squared Error (pixel-level deviation), LPIPS (perceptual discrepancy in deep feature space), SSIM (local luminance, contrast, and texture fidelity), and Local Complexity (stability of intermediate representations under small perturbations). In parallel, non-learned descriptors — Local Binary Patterns for texture and Discrete Wavelet Transform bands for frequency content — are extracted from both the noisy and denoised versions, converted into normalized histograms with K bins and ε-smoothing, and compared using Kullback-Leibler divergence.
The rationale for mixing learned metrics with non-learned ones is that learned feature spaces inherit invariances that can create "blind spots"; texture and frequency statistics supply orthogonal information.
Concatenating all M metrics across all N noise levels gives a vector s(x) ∈ ℝ^{NM}. For tabular data, where perceptual and texture descriptors are not meaningful, the authors use pointwise MSE plus Reconstruction Rank-Order Consistency to check whether pairwise feature orderings are preserved between x and its reconstruction — MSE catches absolute errors, rank consistency catches structural misalignments.
The resulting embedding is evaluated three ways: (1) an Isolation Forest fitted only on in-distribution vectors, giving anomaly scores in the standard unsupervised sense; (2) K-means with the number of clusters fixed to the number of OOD categories; (3) a two-layer MLP trained on held-out in-distribution and OOD samples. The authors stress that they are not claiming their metric choice is optimal for all datasets — the choice should be guided by which shifts matter for a given application.
Why This Matters
Impact on research. The paper challenges what it calls the prevailing paradigm of OOD detection by showing both empirically and theoretically that any single test statistic must collapse heterogeneous shifts into one dimension. The non-identifiability proposition gives a formal reason why "detect the shift" and "understand the shift" are different problems. It also reframes the evaluation question: the paper argues that near-OOD and far-OOD trade-offs, finite-sample effects, and impossibility results in the literature all point to the same limitation.
Real-world applications.
- Medical imaging and sensing: a model that can distinguish "scanner protocol changed" (covariate shift, potentially correctable) from "this is a pathology never seen in training" (semantic shift, needs human review) supports different actions.
- Autonomous driving: sensor degradation, fog, or adversarial stickers on a sign are very different failures from encountering a novel object class; the appropriate response — clean the sensor, slow down, or hand over control — depends on which one occurred.
- Continual and open-ended learning: shifts that signal genuinely new concepts are candidate material for model adaptation, while corrupted or irrelevant data should be discarded. Distinguishing them is a prerequisite for deciding whether to retrain, denoise, or reject.
- Industrial and tabular monitoring: the ADBench experiments suggest the approach transfers to tabular anomaly detection, where a shift may disrupt relationships between features without producing large absolute errors.
Industry relevance. The paper is explicit that DISC's computational cost is higher than single-pass scalar detectors because it requires denoising evaluations at multiple noise levels plus a richer metric suite. The authors position it alongside higher-cost uncertainty estimators such as Deep Ensembles and Monte Carlo Dropout, and argue the overhead is justified when the goal extends beyond binary rejection toward OOD type characterization. That is a straightforward deployment trade-off for teams deciding whether to spend inference compute on richer uncertainty signals.
Future Directions
-
Reducing inference cost. The authors identify computational expense as a natural trade-off of DISC and place it alongside Deep Ensembles and MC Dropout, but do not report a route to cheaper trajectories; compressing the metric suite or the noise-level set is an open engineering question.
-
Extending beyond covariate and semantic shifts. The paper explicitly scopes out label shifts, arguing they primarily reflect downstream curation effects and require different assumptions and estimation tools. Incorporating them would broaden the taxonomy.
-
Principled metric selection. The authors state they do not claim their metrics are optimal for all datasets; a method for choosing statistics based on which shifts matter for a given application remains open.
-
Richer open-world evaluation. The conclusion calls for moving beyond the simple ID-vs-OOD setting toward evaluations where systems reason about the nature of novelty. The multi-OOD clustering and classification protocol introduced here is a step, but the reported accuracies (0.5120 clustering, 0.7131 supervised on CIFAR-10) show substantial headroom.
Target Audience
Researchers and practitioners working on OOD detection, anomaly detection, uncertainty estimation, and generative modeling, particularly those who already use diffusion models and want to extract more information from the denoising trajectory than a single reconstruction error. The multi-OOD evaluation protocol and the "All detectors" concatenation baseline are likely to interest benchmark designers, while the non-identifiability proposition is aimed at readers thinking about the theoretical limits of scalar scores. Readers applying OOD detection in safety-critical or continual-learning settings will find the practical framing — that different shift types warrant different actions — most useful, though they should weigh the reported inference cost.
Authors’ abstract
Detecting out-of-distribution (OOD) data is critical for machine learning, be it for safety reasons or to enable open-ended learning. However, beyond mere detection, choosing an appropriate course of action typically hinges on the type of OOD data encountered. Unfortunately, the latter is generally not distinguished in practice, as modern OOD detection methods collapse distributional shifts into single scalar outlier scores. This work argues that scalar-based methods are thus insufficient for OOD data to be properly contextualized and prospectively exploited, a limitation we overcome with the introduction of DISC: Diffusion-based Statistical Characterization. DISC leverages the iterative denoising process of diffusion models to extract a rich, multi-dimensional feature vector that captures statistical discrepancies across multiple noise levels. Extensive experiments on image and tabular benchmarks show that DISC matches or surpasses state-of-the-art detectors for OOD detection and, crucially, also classifies OOD type, a capability largely absent from prior work. As such, our work enables a shift from simple binary OOD detection to a more granular detection.