Skip to content
AI.info

Generative AI

Hallucination, Uncertainty, Verification, and Abstention

Diagnose unsupported generation and build claim verification, uncertainty handling, abstention, and escalation into the product workflow.

By the end you can

Example

Six opinions, none of them real, and a court that priced the failure

A brief filed in federal court cited six judicial opinions — Varghese, Shaboon, Petersen, Martinez, Durden and Miller — with quotations and citations attached to each. None of the six existed. ChatGPT had produced them. On 22 June 2023 Judge P. Kevin Castel, in the Southern District of New York, issued an Opinion and Order on Sanctions in Mata v. Avianca.

The order opens on the finding: “Peter LoDuca, Steven A. Schwartz and the law firm of Levidow, Levidow & Oberman P.C. (the "Levidow Firm") (collectively, "Respondents") abandoned their responsibilities when they submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question. Many harms flow from the submission of fake opinions.” — P. Kevin Castel, U.S.D.J., Opinion and Order on Sanctions, Mata v. Avianca, 22 June 2023.

The penalty was $5,000, jointly and severally, paid into the Registry of the Court within 14 days. On top of that came letters to the client and to each judge falsely named as the author of a fabricated opinion. That is the shape of the failure this lesson is about. Fluent output, correct form, nothing underneath. And no step in the workflow that asked whether the cited authority was real.

  • Source fidelity: The output carried the exact surface form of law — case names, quotations, citations — and that form was what carried it into a federal filing.
  • Evidence gap: Behind the form there was nothing. Varghese, Shaboon, Petersen, Martinez, Durden and Miller were six non-existent judicial opinions; no retrieval step had ever put a real one in front of the model.
  • Inference leap: The tool was asked for supporting authority and produced authority-shaped text, because producing something is what it does when the evidence is absent.
  • Confidence mismatch: Nothing in the wording marked the boundary between a real docket and an invented one. The respondents stood by the opinions after judicial orders had called their existence into question. The check that should have caught the fabrication confirmed it instead.
  • Product failure: The price was set afterwards, and by a court. $5,000 jointly and severally, paid into the Registry of the Court within 14 days, plus letters to the client and to each judge falsely named as an author.

Comparison

Confabulation is one named risk of twelve, not the whole problem

Different failure types require different corrections, and the vocabulary for saying so does not have to be improvised. NIST published its Generative AI Profile on 26 July 2024. It names twelve generative-AI risk categories, and defines one of them like this: “Confabulation: The production of confidently stated but erroneous or false content (known colloquially as "hallucinations" or "fabrications") by which users may be misled or deceived.” — NIST AI 600-1, Generative Artificial Intelligence Profile. The profile goes further. It notes that confabulated citations and confabulated justifying logic can themselves mislead users into trusting the output. That is exactly the mechanism that carried six invented opinions into Mata v. Avianca.

One named risk out of twelve is a useful reminder. A wrong answer has more than one cause, and the causes take different repairs.

Unsupported generation. The output states a claim absent from the available evidence. Improve retrieval or source coverage, require claim-level support, abstain when evidence is absent. Do not try to fix it with instructions about tone.

Reasoning or calculation error. The evidence is present but the transformation or inference is wrong. Use deterministic tools where possible, decompose and verify intermediate results, add targeted examples and tests. Inspect the trace, not only the citations.

Stale or conflicting evidence. Sources disagree, or no longer match the required time and jurisdiction. Track authority and effective dates, expose conflicts to the workflow, prefer authoritative current sources, escalate unresolved contradictions.

Communication failure. The core answer may be defensible, but its scope or its uncertainty is hidden. State conditions and limitations, distinguish fact from recommendation, use calibrated user messaging, and avoid performative certainty.

FigureComparison · 4 columns

Unsupported generation

The output states a claim absent from the available evidence.

  • Improve retrieval or source coverage
  • Require claim-level support
  • Use abstention when evidence is absent
  • Do not fix solely with tone instructions

Reasoning or calculation error

Evidence is present, but transformation or inference is wrong.

  • Use deterministic tools where possible
  • Decompose and verify intermediate results
  • Add targeted examples and tests
  • Inspect trace rather than only citations

Stale or conflicting evidence

Sources disagree or no longer match the required time and jurisdiction.

  • Track authority and effective dates
  • Expose conflicts to the workflow
  • Prefer authoritative current sources
  • Escalate unresolved contradictions

Communication failure

The core answer may be defensible, but scope or uncertainty is hidden.

  • State conditions and limitations
  • Distinguish fact from recommendation
  • Use calibrated user messaging
  • Avoid performative certainty

Visual

Retrieval is not verification: between 17% and 33%, citations included

It is tempting to treat a citation as the check. Three commercial legal research products were sold on a promise of “hallucination-free” citations, and they were tested against that promise on 202 queries — the first preregistered evaluation of retrieval-augmented commercial legal AI. Magesh, Surani and four colleagues published the result in 2025 in the Journal of Empirical Legal Studies. Their finding: “While hallucinations are reduced relative to general-purpose chatbots (GPT-4), we find that the AI research tools made by LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI) each hallucinate between 17% and 33% of the time.” — Varun Magesh and colleagues, Journal of Empirical Legal Studies, 2025.

Retrieval helped. Retrieval did not finish the job. Up to a third of answers from systems built on real documents still failed. So check at the level where evidence can actually support or refute a statement, not at the level of whether a source is attached.

Segment the output. Identify atomic factual, numerical, procedural, and recommendation claims.

Classify required evidence. Map each claim to a source, tool, policy, calculation, or human authority.

Retrieve or compute. Acquire the evidence using independent components and current permissions.

Compare claim and evidence. Check entailment, contradiction, scope, date, units, and missing premises — the properties a nearby citation does not guarantee.

Revise or abstain. Remove unsupported content, request clarification, or escalate the case.

Preserve the trace. Record claim, evidence, verifier result, and final action for audit and learning.

FigureProcess · 6 steps
  1. 1

    Segment the output

    Identify atomic factual, numerical, procedural, and recommendation claims.

  2. 2

    Classify required evidence

    Map each claim to source, tool, policy, calculation, or human authority.

  3. 3

    Retrieve or compute

    Acquire evidence using independent components and current permissions.

  4. 4

    Compare claim and evidence

    Check entailment, contradiction, scope, date, units, and missing premises.

  5. 5

    Revise or abstain

    Remove unsupported content, request clarification, or escalate the case.

  6. 6

    Preserve the trace

    Record claim, evidence, verifier result, and final action for audit and learning.

Model confidence is not one observable quantity

Token probabilities describe the model’s next-token distribution under a specific context and decoding setup. A verbal phrase such as “I am certain” is another generated sequence, not a calibrated factual guarantee. There is a measured version of that warning. Dahl, Magesh and two colleagues asked models specific, verifiable questions about randomly selected federal court cases, and reported: “Second, we find that legal hallucinations are alarmingly prevalent, occurring between 58% of the time with ChatGPT 4 and 88% with Llama 2, when these models are asked specific, verifiable questions about random federal court cases. Third, we illustrate that LLMs often fail to correct a user's incorrect legal assumptions in a contra-factual question setup. Fourth, we provide evidence that LLMs cannot always predict, or do not always know, when they are producing legal hallucinations.” — Matthew Dahl and colleagues, Large Legal Fictions, 2024. Between 58% and 88%, with GPT 3.5 and then PaLM 2 ranking in between. The system that would have to raise the alarm is the system that does not always know.

The deeper point is that accuracy and calibration move independently, and this was demonstrated long before generative text. A 2017 paper on calibration opens with the result: “We discover that modern neural networks, unlike those from a decade ago, are poorly calibrated.” — Chuan Guo and colleagues, On Calibration of Modern Neural Networks, 2017. Their example is a 110-layer ResNet on CIFAR-100 against a 5-layer LeNet: 30.6% error versus 44.9%. The deeper network is clearly the more accurate of the two, and clearly the more overconfident. Their fix is not a better prompt. It is temperature scaling — a single parameter fitted on a held-out validation set — with miscalibration measured by Expected Calibration Error (ECE), a metric the paper attributes to Naeini et al. (2015).

That is what a validated signal looks like. A held-out set, a metric, and a number you can watch. Reliability can be estimated in many ways: evidence coverage, retrieval scores, verifier agreement, ensemble variation, calibration sets, tool results, historical error rates. Each of them has to be validated for the decision it supports before it is allowed to gate anything, in the way ECE and a calibration set validate a confidence score.

Do not expose an unvalidated internal score as a user-facing probability of truth.

Key idea

Nine benchmarks out of ten pay nothing for admitting ignorance

A system should decline, ask for clarification, or escalate when evidence is missing, the request exceeds its authority, or the expected harm exceeds the value of an automatic answer. The obstacle is not that models are shy about abstaining. Abstention has been trained out of them, and the size of that incentive has been counted. Kalai, Nachum and two colleagues argued in 2025 that hallucination survives post-training because of how systems are graded: “Binary evaluations of language models impose a false right-wrong dichotomy, award no credit to answers that express uncertainty, omit dubious details, or request clarification. Such metrics, including accuracy and pass rate, remain the field's prevailing norm, as argued below. Under binary grading, abstaining is strictly sub-optimal.” — Adam Tauman Kalai and colleagues, Why Language Models Hallucinate, 2025.

Their Table 2 surveys ten leading benchmarks. Nine of them grade strictly correct or incorrect and award no credit at all for “I don't know”: GPQA, MMLU-Pro, IFEval, Omni-MATH, BBH, MATH (L5 split), MuSR, SWE-bench and HLE. The tenth, WildBench, is the only one that gives partial credit, and the authors note that its 1–10 rubric may still score an admission of ignorance below a plausible hallucination. A guess is free. Silence costs.

So abstention is a policy the product has to impose against that gradient. Its settings come from the risk and from how many reviewers there are. If abstention sends every difficult case to an overloaded queue, the product has moved the failure rather than controlled it.

A useful abstention policy balances residual risk, coverage, delay, and downstream review capacity.

Case

817 questions, 58 percent truthful, and humans at 94

The problem is old enough to have benchmarks. TruthfulQA is 817 questions across 38 categories, written so that some humans would answer them falsely because of a false belief or misconception. Lin and two colleagues built it and published it at ACL in 2022. Their headline result is one sentence: “the best model was truthful on 58% of questions, while human performance was 94%”.

Read the two figures the other way round and the gap is starker. On that set the best model answered falsely 42% of the time, and people 6% of the time. Then read the design of the set before generalising. These are questions selected because the false answer is the tempting one. So 94% is human accuracy on a deliberately misleading set, and 58% is measured against exactly that. It is not a general accuracy rate. It is a measurement of what happens where a confident wrong answer is available.

Figure

The truthfulness gap, and the size of the benchmark behind it: 817 questions written so that the confident wrong answer is the tempting one.

Case

238 passages, 1,908 sentences, 27.0 percent of them accurate

A model can also be checked at runtime against itself. SelfCheckGPT samples a model several times and checks the answers against each other, fact-checking a black-box model’s responses “in a zero-resource fashion, i.e. without an external database”, and without access to the output probability distribution. Manakul and two colleagues presented it at EMNLP in 2023.

The evaluation set is worth knowing before the score is. They built it from 238 GPT-3 (text-davinci-003) passages containing 1,908 annotated sentences, and the annotation reads: “Of the 1908 annotated sentences, 761 (39.9%) of the sentences were labelled major-inaccurate, 631 (33.1%) minor-inaccurate, and 516 (27.0%) accurate.” — Potsawee Manakul and colleagues, SelfCheckGPT, EMNLP 2023. Just over a quarter of generated sentences were accurate. That is the pool a consistency score was being asked to sort. It is why a detector that needs no database at all was worth building, and why what it returns has to be read carefully.

Steps

Design verification and abstention together — Article 14(4) requires it

A check is useful only when its result changes what the system does. For high-risk systems in the European Union that is no longer only a design principle. Article 14(4) of the Artificial Intelligence Act, Regulation (EU) 2024/1689, requires that human overseers be enabled to interpret the system’s output correctly, to decide not to use the system or to disregard, override or reverse its output, and to interrupt the system through a 'stop' button or similar procedure. First on the list is this: “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions to be taken by natural persons;” — Article 14(4)(b), Regulation (EU) 2024/1689, Artificial Intelligence Act. The escalation path, the override and the stop have to exist. The reviewer has to be warned about trusting the model. Build the pipeline so that they do.

1. Define claim classes. Separate facts, calculations, policy conclusions, and suggestions.

2. Assign evidence requirements. Specify acceptable sources, tools, dates, and authority for each class.

3. Build measurable signals. Validate retrieval coverage, verifier accuracy, conflicts, and uncertainty indicators — against a held-out set and a stated metric, as ECE is validated against a calibration set.

4. Choose action thresholds. Map signals and harm tiers to answer, qualify, ask, decline, or escalate. The override and the interruption are Article 14(4) obligations, not optional affordances.

5. Test queue effects. Measure abstention rate, reviewer load, delay, and appeal outcomes. A reviewer who is subject to automation bias and buried in a queue is not oversight.

6. Learn from misses. Add incidents and disagreement cases to evaluation without exposing the final test set.

FigureProcess · 6 steps
  1. 1. Define claim classes

    Separate facts, calculations, policy conclusions, and suggestions.

  2. 2. Assign evidence requirements

    Specify acceptable sources, tools, dates, and authority for each class.

  3. 3. Build measurable signals

    Validate retrieval coverage, verifier accuracy, conflicts, and uncertainty indicators.

  4. 4. Choose action thresholds

    Map signals and harm tiers to answer, qualify, ask, decline, or escalate.

  5. 5. Test queue effects

    Measure abstention rate, reviewer load, delay, and appeal outcomes.

  6. 6. Learn from misses

    Add incidents and disagreement cases to evaluation without exposing the final test set.

Position

Self-checking measures agreement, and agreement is not support

Sampling a model several times and comparing the answers needs no evidence from outside the model at all. That is why it is attractive, and it is also where the ceiling is. SelfCheckGPT was built to work — in its authors’ words — “without an external database”: it samples the model repeatedly and checks the answers against each other. What comes back is a consistency score. Consistency is a fact about the samples. Nothing entered the comparison that could make it a fact about the world.

Run it anyway. A model that contradicts itself has told you where to look. But hold it to this lesson’s own rule, that each signal must be validated for the decision it supports, and notice what this one cannot see. A mistake the model makes steadily produces high agreement, and high agreement is the same reading a true claim gives. On the very set the method was measured against, 516 of 1,908 annotated sentences — 27.0% — were accurate. Most of the material a consistency score had to sort was wrong to begin with. Wrong repeatedly is exactly the condition the score cannot distinguish from right.

The adversarial benchmarks say the same thing from the other side. TruthfulQA’s 817 questions were written so that some people would answer them falsely out of a common misconception, and the best model tested was truthful on 58 percent of them, against 94 percent for humans. Those questions were selected to be adversarial, and the figure is not a general accuracy rate. But they are the questions where the wrong answer is the familiar one, and familiarity is precisely what a comparison with the model’s own drafts cannot detect. Six non-existent opinions arrived in a federal filing with quotations attached. Nothing in the model's own drafts would have disagreed about them either.

A check that consults nothing outside the model can only report that the model was consistent.

Reliable generation often means saying less

The strongest system does not maximize the number of questions answered. It maximizes useful, supported outcomes, while exposing uncertainty and routing unsupported cases safely. Every measurement in this lesson points the same way. Retrieval-augmented legal products still hallucinated between 17% and 33% of the time on 202 preregistered queries. Models do not always know when they are hallucinating. Nine of ten surveyed benchmarks pay nothing for saying so. And one afternoon's worth of unverified citations cost $5,000 and a written apology to six judges who had never issued the opinions attributed to them.

The next lesson addresses deliberate attacks and privacy failures. Verification helps with unsupported claims, but adversarial users and untrusted content require additional trust boundaries.

Key takeaways