Research
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study Overview Research area: Natural Language Processing — uncertainty quantification and hallucination det

- arXiv
- 2602.17431
- Published
- 2026-02-19
- Authors
- Dylan Bouchard, Mohit Singh Chauhan, Viren Bajaj, David Skarbrevik
AI summary
Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative StudyOverview
Research area: Natural Language Processing — uncertainty quantification and hallucination detection for large language models, specifically long-form generation.
Technical level: Intermediate. The paper builds a formal taxonomy with mathematical notation but the core ideas (decompose, score, aggregate) are explained in accessible terms.
Scope: This paper organizes, generalizes, and empirically compares black-box methods that assign fine-grained confidence scores to sentences and claims in long-form LLM responses, across five LLMs and two long-form QA datasets.
What This Paper Is About
Most existing methods for detecting hallucinations via uncertainty quantification were built for short outputs — a single answer or span — and they generalize poorly when a model writes several paragraphs containing many separate factual statements. The goal of this paper is to break long responses into smaller units (sentences or claims), score the confidence of each unit, and combine those scores into a response-level confidence, while providing a common framework that lets prior methods be compared directly to one another.
Key Contributions
-
A three-stage taxonomy for fine-grained uncertainty quantification. The authors formalize long-form UQ as a pipeline of (1) response decomposition into sentences or atomic claims, (2) unit-level scoring via a matching scheme, semantic consistency function, and functional form, and (3) response-level aggregation. This provides a common language that unifies prior work and allows apples-to-apples comparisons.
-
Formalization and extension of four scorer families. They define unit-response, matched-unit, unit-QA, and graph-based scorers, showing which configurations reproduce existing methods (LUQ, LUQ-atomic, LUQ-pair, long-form semantic entropy, and the graph centrality method of Jiang et al. (2024)) and which are new generalizations via alternative consistency functions or granularities. They introduce two new graph centralities (Harmonic and Laplacian) and exclude Eigenvector Centrality.
-
A new benchmark: FactScore-STEM-Geo. A 400-question long-form QA dataset spanning four categories across STEM and Geography — Chemical Elements, Scientific Laws, Nerves in the human body, and Mountains. For each category, the 100 entities with the longest Wikipedia articles were selected and paired with prompts instructing the model to write facts about the target topic.
-
A joint empirical comparison within one framework. Rather than comparing methods pairwise, the study evaluates a broad suite of claim-level and sentence-level scorers on five LLMs (Gemini-2.5-Flash, Gemini-2.5-Pro, GPT-4o, GPT-4o-mini, Llama-4-Maverick-17B) and two datasets (FactScore-bio with 500 prompts, and the new FactScore-STEM-Geo with 400 questions). All methods are released in the open-source toolkit
uqlm.
Main Findings
-
Claim-response entailment is hard to beat. Across the ten LLM-dataset combinations, claim-response entailment achieves the highest unit-level AUROC in 3 scenarios and lands within 0.01 of the best scorer in the remaining 7, while attaining the highest AUPRC in all 10. It consistently outperforms more complex claim-level scorers in practice, even when those scorers are more elaborate.
-
Graph-based scorers are close competitors. The top AUROC in the other 7 scenarios comes from graph-based scorers: Closeness Centrality leads in 3, PageRank in 2, and Harmonic Centrality in 1. Among graph-based methods overall, PageRank and Closeness Centrality each lead in three scenarios, with Laplacian and Harmonic each leading in two. Claim-response and graph-based scorers substantially outperform claim-QA approaches.
-
Entailment beats other consistency functions at the claim level. For claim-response scoring, entailment outperforms both contrasted entailment (the function used by Zhang et al. (2024)) and non-contradiction across all ten scenarios.
-
Claim-level scoring generally beats sentence-level scoring. The highest sentence-level AUROC across all scenarios reaches only 0.716 and is consistently lower than claim-level scorers on the same datasets, matching prior observations that hallucination detection is inherently harder at the sentence level. Sentence-level top performers are spread across families: matched-sentence, sentence-response, and sentence-QA each achieve the highest AUROC in 5, 3, and 2 of 10 scenarios respectively.
-
Claim-level AUROC ranges. Scenario-specific top AUROC ranges from 0.67 (Gemini-2.5-Pro responses on FactScore-STEM-Geo with PageRank) to 0.80 (Gemini-2.5-Flash responses on FactScore-Bio with Closeness Centrality). For reference, FactScore-Bio claim-response AUROC by model is 0.794 (Gemini-2.5-Flash), 0.755 (Gemini-2.5-Pro), 0.774 (GPT-4o-Mini), 0.724 (GPT-4o), and 0.791 (Llama-4); on FactScore-STEM-Geo the same scorer yields 0.671, 0.669, 0.718, 0.703, and 0.672.
-
Claim-QA scoring underperforms. Claim-QA scorers rarely achieve AUROC above 0.6. Non-contradiction is the top claim-QA consistency function in 8 of 10 scenarios. The authors note that Farquhar et al. (2024) decompose into coarser factoids containing several atomic claims and achieve better results, and that manual inspection revealed many atomic claims are poorly suited for question-inversion.
-
Verbalized confidence sits in the middle. Unit-level verbalized confidence lags behind graph-based and claim-response scorers but outperforms claim-QA scorers.
-
Calibration is moderate at best for claim-level scorers, which rarely attain ECE below 0.1. PageRank, Laplacian Centrality, and Betweenness Centrality are notable exceptions with much worse calibration, with ECE consistently above 0.6. At the sentence level, Sentence-QA Exact Match yields the lowest ECE in 8 of 10 scenarios. Apart from PageRank, Laplacian Centrality, and Betweenness Centrality, sentence-level scorers tend to be less calibrated than claim-level scorers.
-
Uncertainty-aware decoding (UAD) substantially improves factuality. Filtering low-confidence claims and re-aggregating raises accuracy markedly. For example, filtering Gemini-2.5-Flash responses on FactScore-Bio at the 50th percentile threshold increases accuracy from a baseline of 0.72 to approximately 0.90. Claim-response-based filtering generally yields the most pronounced gains, followed closely by graph-based filtering, with claim-QA-based filtering effective but more modest.
-
Response-level correlations follow unit-level patterns on FactScore-Bio. Closeness and Harmonic Centrality yield the strongest response-level signals, followed closely by claim-response entailment, and these outperform other fine-grained scorers and short-form baselines. Betweenness Centrality exhibits no useful response-level signal. Short-form BERTScore is the most competitive short-form baseline but generally falls short. Short-form white-box scoring (normalized sequence probability) is notably stronger for the GPT models than the Gemini models, with negligible signal for the latter. The authors emphasize that even competitive short-form scorers cannot localize uncertainty in the response or remove low-confidence claims via UAD.
-
Performance drops on FactScore-STEM-Geo. Top response-level correlations by model range from 0.37 to 0.62, compared to 0.60 to 0.74 on FactScore-Bio. Among aggregated claim-level scorers, QA non-contradiction and CR non-contradiction show the highest correlation on this dataset, while the graph-based scorers that dominated FactScore-Bio are rarely competitive at the response level here.
-
More samples help, then stop helping. Ablation studies show performance increases with the number of sampled responses but with substantial diminishing returns, with negligible gains beyond m = 5, consistent with prior work.
Methodology in Plain English
The authors study black-box methods, meaning they only look at the text an LLM produces and never at its internal probabilities. For each prompt, they generate one original response plus 10 sampled responses using stochastic decoding.
Every method follows the same three steps. First, decompose the response into units — sentences found with SpaCy, or atomic claims extracted by an LLM. Second, score each unit by comparing it against evidence drawn from the sampled responses. The comparison uses a semantic consistency function: whether one text entails another (NLI probability), whether it fails to contradict it, a contrasted entailment ratio that excludes the neutral class, normalized cosine similarity, BERTScore, or exact match. Third, aggregate the unit scores into a response-level confidence, primarily by averaging, with minimum, geometric mean, and rank-weighted average also considered.
The four scorer families differ in how the comparison is set up. Unit-response scorers compare each unit directly to full sampled responses. Matched-unit scorers decompose the sampled responses too and compare each original unit to its most similar unit in each sample. Unit-QA scorers turn each unit into a question and measure consistency among the LLM's answers to that question. Graph-based scorers build a bipartite graph of claim-response entailment over the union of unique claims across all responses and score claims by graph centrality — betweenness, closeness, harmonic, Laplacian, or PageRank.
For evaluation, ground-truth factuality labels come from the FactScore grading protocol, which uses an LLM to compare each unit against the corresponding Wikipedia article. At the claim level, claims are further classified as objective or subjective and only objective claims are retained for evaluation; at the sentence level, all sentences are included. Gemini-2.5-Flash handled claim decomposition, claim merging, unit question generation, and grading.
Cost constraints shaped the design: the authors did not compute matched-claim scores because with roughly 25 claims per response and m = 10 samples, that requires about m · N_claim² = 6,250 NLI comparisons per prompt, roughly 25 times the cost of matched-sentence scoring. For unit-QA scoring, they generate an original response and 5 sampled responses per unit question, and use two questions per sentence. NLI scores use microsoft/deberta-large-mnli.
Response-level scoring is evaluated by correlating aggregated confidence with response-level grades, benchmarked against four short-form black-box scorers (BERTScore, cosine similarity, non-contradiction, semantic negentropy) and one short-form white-box scorer (normalized sequence probability). Note that log-probabilities were not exposed by the inference API for Llama-4-17B during the evaluation phase, so results for that model exclude normalized sequence probability.
Why This Matters
Long-form LLM outputs are used in contexts where a single confidently stated but wrong claim can cause real harm, and a response-level confidence number alone cannot tell a reader which sentence to distrust. This work makes fine-grained uncertainty actionable: it shows which scorer to pick, shows that a simple claim-response entailment approach is competitive with or better than more elaborate designs, and demonstrates that filtering low-confidence claims can raise factual accuracy substantially. It also gives researchers a shared vocabulary so future methods can be compared honestly rather than pairwise.
Real-world applications:
- Healthcare and life sciences documentation — flagging which claims in a generated clinical or scientific summary need human verification before use.
- Biography and reference-content generation — automatic quality control for encyclopedic or biographical drafts, where the FactScore-bio and FactScore-STEM-Geo benchmarks apply directly.
- Retrieval-augmented and open-ended question answering — localizing unreliable statements in multi-paragraph answers so downstream systems can drop or re-verify them.
- Selective editing and post-processing — using UAD to retain only high-confidence claims and reconstruct a more factual response before it reaches a user.
Industry relevance: The methods are black-box, so they work with deployed commercial APIs where internal log-probabilities are unavailable — though the paper notes Gemini models exposed so little white-box signal that the white-box baseline was negligible for them. The authors are affiliated with CVS Health, and the work is released in the open-source uqlm toolkit, making it directly usable by practitioners who need deployable reliability tooling rather than only research prototypes.
Future Directions
- Why claim-QA scorers underperform. Manual inspection revealed many atomic claims are poorly suited for question-inversion; determining what makes a claim invertible into a useful question, and whether coarser factoid-level decomposition (as in Farquhar et al. (2024)) closes the gap, remains open.
- Improving calibration. Claim-level scorers rarely attain ECE below 0.1, and several graph centralities (PageRank, Laplacian, Betweenness) exceed 0.6. Methods for calibrating fine-grained scores — and for choosing thresholds — are needed for reliable deployment.
- Closing the gap on harder domains. Performance drops from FactScore-Bio (top correlations 0.60–0.74) to FactScore-STEM-Geo (0.37–0.62), and graph-based scorers that dominated one dataset are rarely competitive on the other. Understanding which domains favor which scorer family would improve practical guidance.
- Extending the taxonomy to more components. Matched-claim scoring was omitted for cost reasons, and the authors note that the decomposition, question-generation, graph-construction, and grading LLMs need not be the same as the generating model — leaving the effect of those model choices, and of aggregation operators beyond averaging, as open design questions.
Target Audience
Researchers working on hallucination detection, LLM reliability, and uncertainty quantification will benefit most, particularly those who need a unifying framework for comparing fine-grained methods. Practitioners building production systems that generate long-form content — and who need black-box, API-compatible confidence signals plus claim-level filtering — will find the practical guidance and the uqlm toolkit directly useful. Readers seeking an accessible entry point to long-form UQ will find the three-stage decomposition intuitive, though comfort with formal notation and standard evaluation metrics (AUROC, AUPRC, ECE, Brier Score) helps.
Note: the provided paper content is truncated at the end of Section 4.4; any discussion, limitations, and conclusion sections are not included in the source text and are therefore not summarized here.
Authors’ abstract
Uncertainty quantification has emerged as an effective approach to closed-book hallucination detection for LLMs, but existing methods are largely designed for short-form outputs and do not generalize well to long-form generation. We introduce a taxonomy for fine-grained uncertainty quantification in long-form LLM outputs that distinguishes methods by design choices at three stages: response decomposition, unit-level scoring, and response-level aggregation. We formalize several families of consistency-based black-box scorers, providing generalizations and extensions of existing methods. In our experiments across multiple LLMs and datasets, we find 1) claim-response entailment consistently performs better or on par with more complex claim-level scorers, 2) claim-level scoring generally yields better results than sentence-level scoring, and 3) uncertainty-aware decoding is highly effective for improving the factuality of long-form outputs. Our framework clarifies relationships between prior methods, enables apples-to-apples comparisons, and provides practical guidance for selecting components for fine-grained UQ.