Research
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Overview Research area: Large language model (LLM) reliability and safety, specifically hallucination detection, reframed as a problem in out-of-distribution (OOD) detection. Technical level: Advanced

- arXiv
- 2602.07253
- Published
- 2026-02-06
- Authors
- Litian Liu, Reza Pourreza, Yubing Jian, Yao Qin, Roland Memisevic
AI summary
Overview
- Research area: Large language model (LLM) reliability and safety, specifically hallucination detection, reframed as a problem in out-of-distribution (OOD) detection.
- Technical level: Advanced. The paper derives geometric scores from linear-classifier theory (decision regions, decision boundaries, Moore–Penrose pseudo-inverses) and assumes familiarity with LLM decoding, AUROC evaluation, and OOD detection literature.
- Scope in one sentence: The paper adapts two existing OOD detectors, NCI (Liu and Qin, 2025) and fDBD (Liu and Qin, 2024), into training-free, single-sample hallucination detectors for LLMs and evaluates them across reasoning benchmarks, model families, and decoding temperatures.
What This Paper Is About
The paper addresses the problem that existing hallucination detectors perform well on short question-answering tasks but struggle on reasoning tasks, where multiple valid reasoning paths make cross-sample consistency checks conceptually difficult. The authors' core idea is that next-token prediction in an LLM can be viewed as a classification problem, so the large body of work on OOD detection — where a classifier is forced to assign a label to an input whose true class was never seen during training — can be repurposed for hallucination detection. Their goal is a detector that requires no training and no multiple samples, works on reasoning tasks, and remains reliable under stochastic decoding.
Key Contributions
-
A conceptual bridge from OOD detection to hallucination detection. The paper formalizes the analogy that both problems reduce to measuring model uncertainty, and then systematically works through the three practical obstacles to transferring OOD methods to LLMs: the infeasibility of estimating training statistics, the dramatically larger label space (vocabulary vs. thousands of classes), and stochastic token sampling.
-
An analytical proxy for the training feature mean. Because computing the mean of penultimate-layer features over an LLM's training corpus is impractical, the authors define a Decision-Neutral Closest Point obtained by minimizing logit variance across the vocabulary, and derive a closed-form solution (Lemma 4.1) that requires no training data.
-
A top-k restriction for the decision-boundary score. To make the fDBD score tractable and less noisy in a vocabulary of hundreds of thousands of tokens, the authors restrict the boundary-distance computation to the k highest-logit alternative tokens, reducing per-step complexity from O(d_model · |V|) to O(d_model · k + |V|).
-
Step-wise scores extended to sequence-level detectors. Both NCI and fDBD are averaged across decoding steps to produce a sequence-level hallucination score, which is then thresholded — giving a single-sample, training-free detector evaluated across multiple benchmarks and models.
Main Findings
-
Hallucinated responses are geometrically distinguishable. Figure 2 (CSQA, Llama-3.2-3B-Instruct) shows features from hallucinated responses have lower proximity to weight vectors (the NCI signal) and smaller distance to decision boundaries (the fDBD signal) than features from correct responses.
-
The analytical proxy outperforms the empirical one. On CSQA with Llama-3.2-3B-Instruct (Table 1), NCI with the analytical proxy reaches AUROC 66.07, versus 62.79 for NCI with an empirical mean estimated from the CSQA training set and 63.23 for the Perplexity baseline.
-
Restricting fDBD to top-k tokens improves performance. On CSQA with Llama-3.2-3B-Instruct (Table 2), AUROC peaks at k = 1,000 with 69.24, compared with 68.15 when all tokens are used. Intermediate values: k=1 → 68.64, k=10 → 68.76, k=100 → 69.18, k=10,000 → 68.87, k=100,000 → 68.26.
-
The detectors survive stochastic decoding. Across temperatures 0.2, 0.5, 0.8, and 1.0 on CSQA with Llama-3.2-3B-Instruct (Table 3, five random seeds), NCI scores 67.07 ± 0.53, 66.04 ± 1.59, 67.53 ± 0.88, and 67.93 ± 1.65; fDBD scores 69.30 ± 0.56, 68.19 ± 1.47, 69.12 ± 0.56, and 69.19 ± 1.98. Perplexity over the same settings scores 63.49 ± 0.75, 62.08 ± 2.34, 63.45 ± 0.71, and 62.68 ± 1.08.
-
Strong results across models and datasets. Table 4(a) reports AUROC on CSQA, GSM8K, and AQuA. For Llama-3.2-3B-Instruct: fDBD with selected k scores 69.24 / 76.36 / 76.20, NCI scores 66.07 / 76.32 / 74.41, and Perplexity scores 63.23 / 69.63 / 72.85. For Qwen-2.5-7B-Instruct: fDBD with selected k scores 72.47 / 77.19 / 78.22, NCI scores 71.60 / 75.83 / 78.19, and Perplexity scores 61.94 / 71.54 / 71.66.
-
Multi-sample baselines underperform on reasoning tasks. Lexical Similarity, SelfCheckGPT-NLI, and Semantic Entropy — which require multiple samples (the experiments use temperature 0.5 and three samples) — are less effective on the reasoning benchmarks than the single-sample OOD-inspired detectors, consistent with the authors' argument that comparing samples is hard when valid reasoning chains diverge.
-
Latency overhead is minimal. On an A100 GPU with CSQA and Llama-3.2-3B-Instruct and k=1000 for fDBD (Table 4(b)), latency in ms/token is 31.94 for standard decoding, 32.88 for Perplexity, 32.54 for NCI, and 32.71 for fDBD.
-
Low sensitivity to the k hyperparameter. Table 5 perturbs the selected k by ±5%, ±10%, and ±20% on Llama-3.2-3B-Instruct. For CSQA, AUROC moves only between 69.22 and 69.25; for GSM8K between 76.32 and 76.41; for AQuA between 76.15 and 76.34.
-
Generalization to open-ended tasks. On TruthfulQA with Llama-3.2-3B-Instruct (Table 6), NCI scores 68.17 and fDBD 67.83, versus 63.78 for Perplexity, 60.62 for Max P, 55.16 for CoE-C, 51.32 for Predictive Probability, 50.38 for LN Predictive Probability, and 48.91 for CoE-R.
-
The detectors trade efficiency for accuracy differently. NCI costs O(d_model) operations per step and is cheaper; fDBD costs O(d_model · |V|) in full form but O(d_model · k + |V|) with the top-k restriction, and in most reported cases fDBD outperforms NCI.
Methodology in Plain English
The authors treat the final linear layer of an LLM as a classifier over the vocabulary: at each decoding step, the penultimate-layer feature is scored against every token's weight vector, and the highest-scoring token is the prediction. This lets them import two geometric uncertainty measures from OOD detection.
The first, NCI, asks how aligned the current feature is with the weight vector of the most confident token, after subtracting an estimate of the mean training feature. The second, fDBD, asks how far the feature is from the boundary at which the model's top choice would flip to a different token. To make fDBD tractable, they use a closed-form lower bound on that distance (Theorem 3.3) that avoids solving an optimization per token.
Three adaptation problems are addressed in turn. First, instead of estimating the mean training feature from data, they solve analytically for the point that minimizes logit variance across the vocabulary — a "decision-neutral" point — and use that as a stand-in. (For Llama-3.2-3B-Instruct, which has a zero-bias language head, this reduces to the origin.) Second, instead of averaging boundary distances over the entire vocabulary, they keep only the top k most likely alternative tokens, chosen with a Quickselect-style procedure. Third, they test robustness to stochastic decoding by sampling at several temperatures with multiple seeds and by checking an alternative variant in which the score is computed relative to the actually decoded token rather than the highest-logit token.
Each per-step score is averaged across the response's tokens to produce a sequence-level score, which is thresholded to yield a binary hallucination decision. Evaluation uses AUROC, a threshold-free metric, and the paper follows prior work in allowing AUROC values outside the conventional [50, 100] range so that methods can be compared even when hallucinated samples receive higher confidence scores than correct ones.
Datasets used are CSQA (commonsense reasoning, multiple choice), GSM8K (free-form numerical answers), AQuA (multiple-choice math), and TruthfulQA (open-ended). The paper reports CSQA validation-set sizes of 1,319 questions in Section 4.1 and 1,221 questions in Section 5.1; GSM8K test set is 1,319 questions and AQuA validation split is 254 questions. Additional generalization results — to Qwen-3-32B, other model sizes in the same family, non-instruction-tuned base models, MoE models, and alternative architectures — are referenced to appendices A through E, but the appendix content itself is not included in the provided text.
Why This Matters
Impact on research. The paper proposes a new framing for hallucination detection that connects two previously separate literatures. If the analogy holds broadly, decades of OOD detection research become directly applicable to LLM safety, and the paper's specific workarounds (analytical training-statistic proxy, top-k boundary restriction,
Authors’ abstract
Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety.