Skip to content
AI.info

Research

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes and What Recovers It

Overview Research area: Clinical natural language processing and automated evaluation — specifically LLM-as-judge methods applied to AI-drafted clinical notes, sitting at the intersection of absence d

arXiv
2608.31016
Published
2026-08-31
Authors
Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris

AI summary

Overview

Research area: Clinical natural language processing and automated evaluation — specifically LLM-as-judge methods applied to AI-drafted clinical notes, sitting at the intersection of absence detection in language models, summarisation factuality, and scribe audit literature.

Technical level: Advanced. The paper assumes familiarity with paired discrimination, AUC, McNemar tests, bootstrap resampling, Cohen's kappa, and prompt-optimisation tooling, though its central argument is stated in plain terms.

Scope: The paper builds a controlled benchmark of clinical notes with certain, graded omissions and uses it to test whether LLM judges can detect missing information, finding that they cannot unless the task is restructured from an open absence question into closed per-fact presence checks.

What This Paper Is About

Ambient AI scribes draft clinical notes at scale, and published human audits find that the dominant error class is omission — information the encounter established that the note fails to record. The standard quality check is an LLM judge, a second model that reads the transcript and the note and flags problems, and this paper asks whether those judges can detect omissions at all. Because no public corpus supplies a trustworthy answer key for absence, the authors build their own benchmark in which a named fact is certainly removed, then test eighteen judge configurations against it.

Key Contributions

  1. A released benchmark where absence is certain and graded. 500 single-error note pairs built from transcript-derived, audited fact sheets: 298 in which a named fact's primary statement is certainly removed, graded by severity and by how much of the fact survives elsewhere, against 202 added-or-altered controls. The evaluation set is 495 pairs over 112 consultations (293 omissions, 202 commissions, 112 clean twins).

  2. The presence/absence asymmetry, isolated by ablation. Eight judge designs crossing three choices (what to check, how to answer, how many times to ask) run at 0.79–0.94 paired discrimination on added or altered content against 0.50–0.63 on omissions, with the asymmetry also reproduced across eight reference judges and a second model family.

  3. A null across five remedy families plus a localisation of the mechanism. Wording changes, criteria, answer format, voting, and GEPA prompt optimisation each move the judge's operating point without creating usable single-note detection; a list of concrete facts to check accounts for about a third of the eventual recovery and closed per-fact verdicts for the rest.

  4. Two working detectors compared head to head. A per-fact enumeration pipeline whose flags name the missing fact and its severity at 2.7% false alarms, against a single GEPA-evolved call that detects more notes (36.9% against 24.6%, p=0.002) at 6.2% false alarms and roughly a tenth of the measured cost per note at one note per consultation.

Main Findings

  • Commissions are detected, omissions are not. Across the eight ablation designs, paired discrimination runs 0.792 to 0.944 on added or altered content, six of them at 0.87 or above, while the same eight read 0.500 to 0.634 on omissions. The faithfulness-only yes/no judge sits at exactly a coin flip (0.500) with the clean twin in hand.

  • The best monolithic judge is still weak and calibration is not the limit. Faithfulness plus completeness, scored, eight samples reaches only 0.634 (0.601 to 0.655 across replicates) on omissions. Threshold-free, area under the curve stays between 0.503 and 0.575 across all eight designs, and swept to a 10% false-alarm ceiling the best scored design reaches 8.3% detection at 6.5% false alarms.

  • Restatement is the hardest case. Within the best design, complete removals score 0.690 and fragment traces 0.607, but restatement traces — where a second mention of the fact survives elsewhere — collapse to 0.526, a coin flip. By severity, critical reads 0.683, supporting 0.586, peripheral 0.568.

  • No design is a usable single-note detector. The two yes/no completeness designs flag 40–42% of omission notes against 32% of clean notes, but on a per-note majority that gap is 1.2 to 1.7 standard errors, and a judge that flags a third of perfect notes is not deployable. Among scored designs the largest gap is 1.02 standard errors; two of the eight run marginally negative.

  • Reference judges reproduce the pattern. The deployed faithfulness judge scores 0.518 on omissions against 0.886 on commissions, G-Eval 0.568 against 0.890, the checklist judge 0.647 against 0.744, the engineered judge 0.639 against 0.754, and the first optimisation campaign's winning prompt 0.549 against 0.861. All 21 candidate prompts in that campaign detect commissions better than omissions, by 14.6 to 49.3 percentage points.

  • The asymmetry survives a change of model family, with one exception. Re-running two designs with a Gemini-family judge gives commissions at 0.951 and 0.961 against omissions at 0.551 and 0.683. That family's completeness-scored design does clear its own noise on single notes in all three replicates (19.1 to 20.5% detection at 0.9 to 3.6% false alarms), which no primary-family design did, so single-note claims are scoped to the family measured.

  • The commission arm holds on independent physician labels. On a frozen 300-text stratified sample of MEDEC (156 errored, 144 error-free), the best of the eight reaches paired 0.827 and AUC 0.811, with 51.3% of errored texts detected at a 10% false-alarm rate, against AUC 0.575 on the paper's own omissions.

  • Remedies shift the operating point without separating. Adding an "AND complete" clause on the MEDEC labels raised detection from 87.8% to 96.8% (p=1.2e-4) and false alarms from 52.8% to 72.9% (p=4.9e-6), while AUC moved from 0.811 to 0.800 and detection at a matched 10% false-alarm rate was identical to three decimal places.

  • Human reviewers do not share the asymmetry. In a planted-error study of AI-drafted patient messages, reviewing physicians caught omissions at roughly the same rate as objective errors, about a quarter to a third of each — a different task on a different corpus, so indicative only.

  • Restructuring the task recovers detection. Converting "is anything missing?" into "list the facts the transcript establishes, then verify each against the note" is arrived at independently by a per-fact pipeline and by a GEPA-evolved single call. The pipeline names the missing fact and its severity at 2.7% false alarms; the single call detects 36.9% of notes against 24.6% (p=0.002, consultations as the unit) at 6.2% false alarms, at roughly a tenth of the cost per note or a third once per-consultation stages are amortised.

  • Human adjudication sides with the pipeline. A physician author validated 70 items across six stages, three of them structurally blinded, and on the ten notes where the pipeline and the best monolithic judge reached opposite conclusions sided with the pipeline on all ten (p=0.002).

  • The severity rubric reproduces across clinicians. Unanchored, two frontier models agreed at Cohen's kappa 0.177 [0.04, 0.31]; with the written rubric, two cross-family graders reached 0.662 [0.59, 0.73] on 683 graded traps. An independent clinician graded the rubric blind across two census sittings with 9 of 12 grades exact and 3 one grade apart; across both sittings 25 of 32 grades were exact with no disagreement exceeding a single grade.

  • Reference notes cannot serve as an answer key. All 53 PriMock notes (100%, interval [93.2, 100]) and all 45 ACI-Bench notes (100%, [92.1, 100]) carry at least one material discrepancy with their own transcript, at means of 10.7 and 7.1 discrepancies per note, with 345 of PriMock's 1,030 graded critical.

  • Real vendor notes break threshold transfer. On 261 of the companion census's 565 notes (87 carrying a panel-verified omission, 174 with no verified finding), no benchmark threshold transfers, but the single call re-calibrated still detects more than the best of the eight at half its false-alarm rate, and three quarters of the notes the pipeline flags name the very fact the census's panel verified.

  • The defeat case is named. Omissions whose fact survives as a restatement elsewhere in the note defeat both recovered routes and are left as an open problem.

Methodology in Plain English

The authors could not borrow an answer key, because they first audited public clinician reference notes against their own transcripts and found every one of them materially discrepant. So they built the key from the transcript instead: extract the facts a consultation establishes, audit that fact sheet with model critics, repair and verify the reference note against it, then mechanically remove one fact per pair. Two independent checks guard the build — blind recovery against 650 facts human authors had written down (646 recovered, 99.4%) and a cross-family panel that verified 71 of 126 complete-removal attempts and 114 of 151 partial removals that leave a trace by design.

Each removed fact is graded on a written rubric by two models from different families, taking the lower grade where they disagree, and labelled by whether nothing, a fragment, or a full restatement of the fact survives elsewhere in the note. Detection is scored two ways: paired discrimination, which asks whether the flawed note scores below its clean twin and is therefore a ceiling production never has, and single-note detection, which reports a judge's flag rate on flawed notes beside its false-alarm rate on clean notes. A judge's thresholds and flag rules were fixed before every run, repeats collapse to a per-note majority, and resampling is at the consultation rather than the note. Human validation runs alongside through blinded physician review and an independent clinician grading the rubric.

Why This Matters

Impact on research. The paper supplies the benchmark the omission literature has lacked and shows that the standard evaluation instrument for deployed scribes is blind to the error class human audits rank first — a failure the authors trace to the open absence question giving the model nothing concrete to check. It also shows that a widely assumed fix, better prompting, moves the operating point rather than the underlying separability, and that reporting a flag rate without the matching false-alarm rate hides the difference.

Real-world applications:

  • Vendor and buyer quality assurance for ambient scribes, where a per-fact check can name the missing fact and its severity rather than returning an unactionable flag.
  • Regulatory and procurement review of claims that an automated quality layer catches documentation errors.
  • Clinical documentation improvement programmes that need to know whether a note recorded what the encounter established.
  • Deployment of cheaper single-call judges calibrated to be more sensitive than the monolithic baselines at a comparable false-alarm rate.

Industry relevance. The cost comparison is framed for production: the single-call judge runs at roughly a tenth of the measured cost per note at one note per consultation, a third once per-consultation stages are amortised, which is the difference between an affordable and an unaffordable evaluation layer at scribe scale.

Future Directions

  • Solving the restatement case. Omissions whose fact survives as a restatement elsewhere in the note defeat both the pipeline and the single call, and the paper names this as its open problem rather than claiming to solve it.
  • Extending single-note detection beyond one model family. The Gemini-family completeness-scored design cleared its own noise where no primary-family design did, but since that endpoint refuses a no-reasoning setting, family and reasoning budget move together and the two cannot currently be separated.
  • Establishing an external anchor for the omission arm. The commission side is anchored on physician-labelled MEDEC texts; the omission side has no equivalent external check, for reasons the paper gives in its discussion section.
  • Making benchmark operating points transfer to real vendor notes. No benchmark threshold transferred to the census notes, requiring re-calibration in the one case that was tested.

Target Audience

Researchers and practitioners working on clinical NLP, LLM-as-judge evaluation, and summarisation factuality; clinical informatics and quality teams evaluating or procuring ambient scribes; and evaluation engineers who need to understand why an omission-detection prompt moves its operating point without becoming a usable detector. The paper is also relevant to anyone designing benchmarks for absence or completeness, since its construction protocol, rubric, and the audit showing that clinician reference notes carry 7 to 11 material errors per note are reusable independently of the judge findings.

Authors’ abstract

Ambient AI scribes draft clinical notes, and published audits find their dominant error is omission: information the encounter established that the note fails to record. The standard check is an LLM judge: a second model reads the note against the transcript and flags problems. We ask whether judges detect omissions. Public corpora cannot supply the answer key: their clinician reference notes and transcripts are materially discrepant. Our benchmark has 500 single-error note pairs from audited fact sheets, 298 with a named fact certainly absent and 202 added-or-altered controls. Across eight judge designs, paired discrimination (the flawed note below its clean twin, 0.5 a coin flip) reads 0.79-0.94 on added or altered content and 0.50-0.63 on omissions. On single notes, no design flags omissions reliably more often than perfect notes. Wording changes, voting and GEPA prompt optimisation move the operating point without creating usable detection. Restructuring the task recovers it: list the facts the transcript establishes, then check the note for each. Two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call. The pipeline's flags name the missing fact and its severity at 2.7% false alarms. The single call detects more (36.9% against 24.6%, p=0.002) at 6.2% false alarms and a tenth of the cost per note. A physician author validated 70 items and, where the two routes disagree, sided with the pipeline on 10 of 10 (p=0.002). A second clinician, not an author, graded the severity rubric blind and agrees to within a grade. On real vendor notes from a companion census no benchmark threshold transfers, but the re-calibrated single call detects more than the best of the eight at half its false-alarm rate. Omissions whose fact is restated elsewhere defeat both routes. We release the benchmark, prompts and judgements.

Read the original paper