Research
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Overview Research area: Evaluation methodology for large language models used in patient-facing medical consultation (clinical NLP / health AI safety). Technical level: Intermediate. The clinical reas
- arXiv
- 2608.17330
- Published
- 2026-08-18
- Authors
- Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang
AI summary
Overview
Research area: Evaluation methodology for large language models used in patient-facing medical consultation (clinical NLP / health AI safety). Technical level: Intermediate. The clinical reasoning and evaluation design are accessible; some familiarity with LLM prompting and benchmarking helps. Scope: A structured, small-sample demonstration that current LLM evaluations skip the "preformulation" stage — the first-contact work of turning a vague patient concern into a clinically usable problem — and that this stage can be measured directly through observable interaction behaviors.
What This Paper Is About
Real medical consultations often begin with a vague, misspelled, minimized, or misframed concern ("throwing up feel sick maybe food poisning"), yet published evaluations of medical LLMs typically start from cases that investigators have already made clinically clear. The authors argue this mismatch creates a preformulation gap: benchmark scores earned on already-formulated cases are used to support safety claims about first contact, a stage the evaluation never tested. Their goal is to translate that gap into observable behaviors — whether a model asks decision-relevant questions before giving advice, corrects unsafe patient plans, calibrates urgency under uncertainty, and leaves a handoff record that can enter supervised care.
Key Contributions
- Names and defines the preformulation gap — the distance between a raw, unstructured patient concern and a clinically usable problem formulation — and argues it is currently inferred from diagnostic accuracy or final-answer quality rather than measured.
- Builds a behavioral test harness using four physician-authored multi-turn vignettes that open with benign, misspelled lay messages and disclose risk-relevant facts over later turns, across three API models and two conditions.
- Combines fixed scripts with adaptive standardized-patient simulation, so that fixed scripts guarantee all models see the same information while adaptive runs test whether models actively elicit decisive facts rather than passively receiving them.
- Specifies a function-level scoring scheme (usable concern, premise repair, safe routing, handoff readiness) and reports observable marker counts and excerpts rather than inferential statistics.
Main Findings
- Advice before elicitation was the dominant baseline behavior. Self-care or home-management advice appeared before any patient answer in 9 of 12 baseline case-model cells, versus 0 of 12 in the instruction condition. Several baseline replies both advised and asked questions in the same message, so the instruction separated elicitation from advice rather than merely adding questions.
- A short workflow instruction changed sequencing and documentation. Structured handoff summaries appeared in 0 of 12 baseline cells and 10 of 12 instruction cells. First replies asking three or more urgency-relevant questions rose from 8 of 12 to 12 of 12.
- Unsafe-plan correction was mostly present in both conditions. Explicitly stated unsafe plans were corrected in 8 of 9 relevant baseline cells and 9 of 9 instruction cells, so the main difference was sequence and directness, not whether correction happened at all. The one baseline miss involved the OpenAI model reading "im foing to drink and sleep" as referring to water; under instruction the same model stated plainly, "I would not recommend drinking alcohol and going to sleep."
- Alarming openers triggered default safety behavior; benign openers did not. The chest-pain case was the exception — all three baseline models opened with emergency red-flag lists and questions, and the gap between conditions was smallest. The sequencing difference was most visible when the opener sounded benign.
- Improved sequencing did not fix elicitation. At the first vomiting turn, chronic illness or medication use was asked in 0 of 3 baseline model runs and 1 of 3 instruction runs. In the adaptive runs, diabetes surfaced before the prespecified disclosure in 1 of 3 baseline vomiting runs and 0 of 3 instruction runs, while the preceding long car ride surfaced in all six calf-pain runs.
- Miscalibration could run in both directions. One baseline model escalated to emergency-level heart-attack language in the upper abdominal pain case after age was disclosed but while pain location and severity were still unreported, later calling the presentation "a medical emergency until proven otherwise." Other baseline replies used relatively categorical clot language without chest symptoms.
- Models showed the relevant knowledge once decisive facts appeared. After scripted disclosures, all three models recognized possible diabetic ketoacidosis or venous thrombosis and escalated.
- Handoff presence was prompt compliance, not quality. Because the instruction explicitly requested a handoff, the authors note this difference primarily demonstrates prompt compliance; summaries were not independently rated for usefulness, completeness, or safety.
- Cited context from other work: in a randomized study of 1,298 members of the public, LLMs tested alone identified the correct condition in 94.9% of scenarios but participants using the same models identified it in fewer than 34.5% of cases (Bean et al. 2026); in a ChatGPT Health evaluation, more than half of gold-standard emergencies were undertriaged to evaluation within 24–48 hours, and minimizing by family or friends shifted triage significantly in edge cases (odds ratio 11.7, 95% confidence interval 3.7–36.6), mostly toward less urgent care (Ramaswamy et al. 2026); and in a classic outpatient study, the diagnosis reached after reviewing the referral letter and taking the history agreed with the diagnosis ultimately accepted in 66 of 80 new patients (Hampton et al. 1975).
Methodology in Plain English
The researchers wrote four vignettes, each a two- to four-turn script in everyday language with deliberate misspellings. Every script starts with a message that could plausibly invite benign self-care advice, then reveals risk over later turns. The four cases covered an older adult with upper abdominal pain, a young adult with pleuritic chest pain, vomiting with diabetes risk, and calf pain after travel.
On July 28, 2026, they ran three API models chosen to approximate consumer assistants — OpenAI chat-latest (ChatGPT), Google gemini-3.5-flash (Gemini), and Anthropic claude-sonnet-4-6 (Claude). Each model received each case twice. In the baseline condition it saw only the patient turns, with no system prompt, no tools, and no indication it was being evaluated. In the instruction condition, the same turns were preceded by a short entry-to-care instruction that added no case-specific diagnoses or medical facts — only workflow sequencing: make the concern clinically usable first, catch unsafe plans, route to the appropriate level of care, and hand off cleanly. Gemini and Claude used temperature 0.2 and a 4,096-token response cap; the OpenAI alias rejected temperature override and used its default.
Because a fixed script hands over decisive facts on schedule whether or not the model asks, the vomiting and calf-pain cases were also run adaptively. An LLM patient simulator (gemini-3.5-flash, 80-token cap to keep replies terse and text-message-like) answered only the questions the tested model actually asked across four patient turns, with prespecified statements about unsafe intended actions inserted at the same points across runs. This produced 24 fixed-script transcripts and 12 adaptive transcripts. Each transcript was reviewed against four prespecified domains — usable concern, premise repair, safe routing, handoff readiness — by counting observable markers. The authors report marker counts and example excerpts; they report no inferential statistics.
Why This Matters
The paper reframes what a medical-consultation benchmark should measure: not answer quality after a case has been given, but interaction quality before the case has been formed. It argues that strong diagnostic accuracy should not be inherited as evidence of triage or first-contact safety, because these tasks differ in what the patient contributes, what the system must elicit, and what harms follow from interaction failure. It also notes that the failure it describes is cheap to influence — a short system instruction changed observed behavior — but fragile, since prompt effects may shift with model updates and did not guarantee elicitation coverage.
Real-world applications:
- Benchmark design for patient-facing LLMs: start items from raw patient openings, preserve misspellings, minimization, mistaken attribution, unclear timelines, and unsafe plans, and combine fixed scripts with adaptive simulation.
- Product specifications for first-contact health assistants: state which minimum information is asked first, how unsafe plans are interrupted, when same-day versus emergency care is recommended, when the system abstains, and how a handoff is generated.
- Generative AI wellness apps that are not marketed or regulated as medical tools but still receive clinically relevant concerns (De Freitas and Cohen 2024).
- Evaluation reporting standards: reports should specify model version, interface, date, system prompt or custom instruction, tool use, and how the patient script advanced, so readers can tell whether a result reflects model capability, product default policy, or the task frame supplied by the experimenter.
Industry relevance: The findings speak directly to model vendors, consumer health product teams, and health systems running triage lines. The authors note that instruction-level fixes should be treated as stopgaps requiring their own evaluation, not substitutes for training and benchmark design that target the preformulation stage.
Future Directions
- Ablate the instruction. The four steps were not ablated, so it is unknown which component carries the main effect.
- Independently rate handoff usefulness. Handoff presence was a binary marker, not a blinded rating of clinical quality, completeness, or safety.
- Scale beyond the demonstration. The four English vignettes, with one run per model and condition, cannot estimate prevalence or rank vendors; the 12 case-model cells are repeated evaluations rather than independent clinical observations.
- Add comparators and controls. The study included no human comparator and no benign control set, and its routing expectations and transcript reviews were defined by the study team without independent blinding or external validation. The adaptive patient simulator was itself an LLM and may differ from real patient answers, and the tested API models approximate but cannot reproduce consumer web products.
Target Audience
Clinical AI and health-NLP researchers designing evaluations of patient-facing models; benchmark and evaluation-methodology groups; product and safety teams building consumer health assistants or wellness apps that receive clinical concerns; clinicians and health-system leaders who interpret vendor claims about diagnostic performance; and regulators or reviewers assessing what a reported benchmark score does and does not support.
Authors’ abstract
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.