Skip to content
AI.info

Research

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

Overview Research area: Natural Language Processing / clinical documentation AI (ambient AI scribes), with a strong measurement-methodology component. Technical level: Advanced. The paper is readable,

arXiv
2608.31017
Published
2026-08-31
Authors
Sebastian Fox, Luke Markham, Ryan Lail, Michael Karotsieris

AI summary

Overview

Research area: Natural Language Processing / clinical documentation AI (ambient AI scribes), with a strong measurement-methodology component.

Technical level: Advanced. The paper is readable, but its core argument concerns measurement instruments, adversarial LLM verification panels, bootstrap intervals over clustered data, and taxonomy construction.

Scope: A verified, evidence-quoted audit of 618 failures produced by three deployed commercial AI scribes across 565 notes written from the same 142 consultations, plus a controlled decomposition of how much the verification instrument itself — not the scribes — moves the resulting failure rate.

What This Paper Is About

Ambient AI scribes draft clinical notes at scale, and the reassurance offered is that a clinician signs every note. The authors wanted to know what those notes actually get wrong, and — separately — how much a "failure rate" depends on the auditing instrument that counted it rather than on the products being audited. They ran three commercial scribe products over the identical set of 142 consultations, adversarially verified every candidate error, built a taxonomy of the survivors, and then deliberately varied the review standard, the reviewing model family and the panel architecture to measure each one's contribution to the headline number.

Key Contributions

  1. A verified, evidence-quoted census and failure taxonomy of three deployed scribe products on a shared corpus, with cross-product replication measured and two clinicians' blinded review of the instrument's judgement layers.
  2. A controlled decomposition of the counting instrument's effect on the count — the review instruction, the model family and the panel architecture measured separately, with which claims survive a strict standard and why, and the instrument released and re-runnable.
  3. Evidence that the divergent error mixes of published scribe audits are consistent with instrument differences of the size measured here.
  4. A practical reading for the clinician who signs and the buyer who compares, released alongside all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.

Main Findings

  • One note in three carries a verified failure. Pooled across products, 177 of 565 notes contain at least one verified failure: 31.3% [27.0, 35.6]. Treating notes as independent instead gives [27.6, 35.3]. The count survived two stages that can only push it down: twelve discovery passes proposed 13,678 candidates, and the verification panel refuted roughly nine in ten of the 5,898 that cleared the importance filter, leaving 618 findings — 10.48% [8.9, 12.1] of candidates put to it.
  • The products differ, and the authors do not read that as a ranking. Scribe A produced 0.454 verified findings per note (60 of 282 notes = 21.3% [16.0, 27.0]), Scribe B 1.525 (57 of 141 = 40.4% [32.6, 48.9]), Scribe C 1.937 (60 of 142 = 42.3% [33.8, 50.0]). Scribe A contributes two notes per consultation through an API at two templates (0.447 and 0.461 findings per note), and capture paths differ — audio replayed for B and C, transcript through an API for A. Under a lenient review standard the gap between products narrows from four-and-a-half-fold to 1.2-fold, and Scribes B and C change order, though Scribe A reads lowest under both standards.
  • Failures concentrate in a few areas. The top three tiers: wrong output 207 (33.5%), addition 181 (29.3%), omission 143 (23.1%), irrelevant or misplaced text 46 (7.4%), and 41 (6.6%) unmapped. The largest discovered subcategory is allergy status and medication list omissions (n=111, about 47 distinct errors, over 33 consultations, 71 rated high importance, findings from all three products), including a paracetamol dose halved in the note ("500mg up to four times daily" against a stated "up to two 500mg paracetamol tablets four times a day"), ibuprofen replaced by nifedipine, and Cetraben becoming cetirizine. The second largest is invented patient identity — name and sex (n=93, about 49 distinct errors, over 34 consultations, 49 high importance), including a patient who gives his name as John Smith documented as "a 32-year-old woman", and an inaudible name filled in as "Gemzar", a chemotherapy brand name.
  • Telephone consultations acquire a physical examination. Remote consult history written as objective exam (n=42) is the fifth-largest cluster: two of the three products produced notes for telephone consultations that document an examination, in the signed section a later reader treats as observed fact. On one such consultation the clinician says out loud that they would normally examine and cannot (Scribe C).
  • One failure mode had no place in the published taxonomies. "Retracted device captured as delivered care" (n=11, 7 from Scribe A and 4 from Scribe B, over a single consultation, 5 high importance): a clinician retracts a thumb spica in the next breath and substitutes a wrist brace, and the note records the spica as applied. The authors name it from a single consultation across two products.
  • Two of the design's most plausible inflations were tested and set aside. No product was given a patient record, demographics or encounter date. Removing invented identities and invented dates while keeping every note in the denominator lowers the headline from 31.3% [27.0, 35.6] to 24.8% [20.8, 29.0], and inverts which product reads worst. The authors' own trap-seeded authored scenarios ran lower (18.3%, 22 of 120) than PriMock57 (43.4%, 99 of 228) and ACI-Bench (28.3%, 51 of 180), so removing everything they wrote would raise the pooled rate, not lower it.
  • The review instruction is the dominant moving part. With the model, evidence and every setting held fixed on a stratified fifth of the notes (1,295 of the 5,898 candidates), the strict instruction verifies 9.3% of candidates and the lenient one 79.0%. The reviewing model family moves the headline too: run alone at the same strict instruction, the gentler family flags 54.8% of sampled notes against the harsher family's 27.8% — roughly double. Adding the gentler model as a second opinion with a tiebreak adds one percentage point. Between 28% and 97% of sampled notes carry a verified failure depending on the standard applied.
  • Human checks upheld the findings but could not separate accuracy from deference. A physician author upheld 20 of 21 randomly drawn verified findings (95.2%, Wilson [77.3, 99.2]), and an independent clinician, not an author and with no study involvement, upheld 12 of 12 (Wilson [75.8, 100]) over 11 consultations — raising the number of census findings a human has assessed from 21 to 33. Both judged every sampled refusal genuine. The authors note the two sittings are never pooled, and that because the physician author judged 29 of 30 items and the independent rater 16 of 16, neither sitting can separate "the flagged findings are nearly all real" from "clinicians asked to check flags accept them."
  • The census disagrees with published audits in a way instrument differences can produce. Omission is 54–86% of errors across the compared audits; here it is 23.1%. The one study that counts over notes as this one does reports omission in 18% of notes, where this census finds 15.4%.
  • Findings are verified findings, not distinct errors. On a model grouping of findings by whether one correction would resolve them together, the 618 verified findings are approximately 265 distinct errors; a model-free grouping by description similarity reads 352 to 529 over the same findings, so the deduplication factor is between about 1.2 and 2.3. The 17 discovered subcategories cover 563 of the 618 findings, with 55 unassigned.

Methodology in Plain English

The authors assembled 142 consultations — 57 recorded UK primary-care consultations (PriMock57), 45 US ambulatory encounters (ACI-Bench), and 40 scenarios they wrote themselves, 30 seeded with documentation traps and 10 written without knowledge of the trap scheme. They replayed or fed each consultation to three commercial scribe products, producing 565 notes (282 from Scribe A, 141 from Scribe B, 142 from Scribe C). No real patient appears anywhere in the study, and no product received a patient record.

Discovery was deliberately over-inclusive: eleven targeted passes, each hunting one known failure mode across every note, plus a twelfth open pass for anything unforeseen, all run on anthropic/claude-opus-5. These proposed 13,678 candidates, of which 5,898 cleared the discovering model's own importance filter. Verification was deliberately hostile: two skeptics from different model families — anthropic/claude-opus-5 (the harsher) and openai/gpt-5.5 (the gentler) — each read the full note and full transcript and were told to refute a candidate if refuting it was at all defensible; a failed or unparseable reply counted as a refutation. A finding counted only if both kept it, or if a third model, openai/gpt-5.4 at high reasoning effort, settled a split upward.

To measure the instrument itself, the authors re-reviewed a stratified fifth of the notes with the model, evidence and every setting unchanged, swapping only the strict instruction for a lenient one. They then altered the panel architecture separately. Products were compared on identical consultations, and the whole pipeline — prompts, model versions, verdicts, findings with evidence quotes — is released so anyone can re-run it on other products.

Why This Matters

Impact on research. The paper argues that a failure rate is a joint property of the scribes and the instrument that counted them, and it supplies the first controlled measurement of that instrument's share. Because none of the human-review audits it compares against reports what its own review standard contributed, the wide divergence in published omission rates (54–86% of errors) cannot currently be attributed between scribe and standard. This census offers a size for the instrument effect, and it matches the one comparison study that counts the way it does.

Real-world applications:

  • Clinician sign-off. The authors' practical reading is that a signing clinician stands where attention matters most, and the taxonomy names what to look for — allergy and medication state, invented identity, and examinations documented for consultations that could not contain one.
  • Procurement and product comparison. A buyer comparing scribes now has a sharper question to ask: an error rate under what instrument? Products here look four-and-a-half times apart under a strict standard and 1.2-fold apart under a lenient one.
  • Re-runnable auditing. The released pipeline, prompts, model versions and 618 evidence-quoted findings let any other product be put through the same published standard, on identical consultations.
  • Rubric calibration. A blinded rubric exercise used facts with deliberately missing content at known severity grades, giving 70% exact agreement with every disagreement within one grade, and a separately regraded set of upheld findings (16 exact, 4 graded above the rubric, none below).

Industry relevance. Vendor self-evaluations are increasingly instrument-aware but are not independent, and the paper notes that a favourable finding

Authors’ abstract

Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.

Read the original paper