Skip to content
AI.info

Research

Towards a Deterministic Math Solver for Clinical Language Models

Overview Research area: Clinical artificial intelligence — specifically making large language models reliable at executing clinical calculators (formulas that convert patient variables into scores, do

Towards a Deterministic Math Solver for Clinical Language Models
arXiv
2609.10728
Published
2026-09-09
Authors
Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi

AI summary

Overview

Research area: Clinical artificial intelligence — specifically making large language models reliable at executing clinical calculators (formulas that convert patient variables into scores, doses, dates or risk estimates).

Technical level: Intermediate. The paper assumes familiarity with LLM benchmarking, clinical formulas and basic statistics, but the core idea — let the model write code instead of doing arithmetic — is easy to grasp.

Scope in one sentence: A controlled, five-seed comparison on 1,100 MedCalc-Bench Verified cases of case-specific program generation executed by a restricted local executor against direct model arithmetic and a hand-written 22-calculator library, using open-weight Qwen2.5 models plus Mistral-7B and Phi-3.5-mini checkpoints, after auditing the benchmark's own formulas against current guidelines.

What This Paper Is About

Clinical calculators such as APACHE II or Cockcroft-Gault must select a formula, pull variables out of a free-text note and do the arithmetic — and LLMs make numerical errors that change the recommendation. The usual fix is to hand-code and validate each calculator as a function, one at a time, which leaves every calculator outside that set unanswerable. The paper tests an alternative interface the authors call Program-Solve: the model does not calculate at all, it writes case-specific Python that a restricted local executor runs as a deterministic solver, so the model's job reduces to deciding how to use the formula it is given.

Key Contributions

  1. A program-generation interface with no calculator-specific code. Program-Solve supplies the model with the clinical formula text and gold variables; the model writes a short Python program, and a restricted executor with no calculator-specific functions or constants runs it and returns a number, date or gestational-age tuple. It supports all 55 calculators and attempts every case, unlike a partial library.

  2. A matched comparison design. Program-Solve is compared against Open-book arithmetic under matched formula access, variable access and note length (both read the whole note), with paired gaps reported under 95% calculator-cluster bootstrap intervals (10,000 draws, seeds kept together). Secondary tests are a case bootstrap, exact McNemar and a sign-flip permutation p, Holm-corrected within family.

  3. An audit of the benchmark itself. The authors audited all 55 supplied formula texts and flag 16 of 55 with a version, use or coefficient concern — four of them replaced by a current guideline (Framingham hard CHD, MDRD GFR, MELD Na, SIRS) that the benchmark still rewards reproducing exactly. A separate completeness audit found 28 of 55 supplied texts wrong or incomplete, 10 of which could not reproduce the benchmark's own number.

  4. Partial-library and hybrid baselines. The hand-written library covers 22 laboratory, physical and date calculators (440 of 1,100 cases) and abstains elsewhere; the paper quantifies the coverage-versus-accuracy trade-off and evaluates library-first hybrids with program or arithmetic fallbacks.

Main Findings

  • Program-Solve is not a reliable advantage at 7B but is one at 32B. With formula, gold variables and whole-note access matched, Qwen2.5-7B reaches 75.31% with the solver against 72.02% for Open-book arithmetic, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38] that crosses zero. Qwen2.5-32B-AWQ reaches 90.53% against 83.47%, +7.05 with an interval of [0.47, 14.60], clear of zero.

  • Failure modes differ between the two routes. Program-Solve returns no valid answer on 6.7% and 0.7% of case-seed rows and a wrong answer on 18.0% and 8.8%, against none unanswered and 28.0% and 16.5% wrong for Open-book arithmetic.

  • The library's advantage is coverage, not execution. The 22-calculator library is exact on its 440 supported cases (100.00%) but abstains on the other 660, giving 40.00% full-set accuracy by construction. Against it, Program-Solve leads by +35.31 pp at 7B and +50.53 pp at 32B on the full set, entirely because the library abstains. On the 440 cases both cover, the library is more accurate: 100% against 84.20% and 98.64% for Program-Solve.

  • Removing formula and variable access is costly. Blind Program-Solve, given neither the formula nor the gold variables, reaches 28.04% at 7B and 44.71% at 32B — declines of 47.27 and 45.82 pp. Against the Extract-Solve library the point estimate is lower at 7B (−10.80) and higher at 32B (+5.71), but both intervals cross zero. The design does not separate recall from extraction.

  • Hybrids beat every single arm. A Gold-first hybrid (library on its 440 cases, Program-Solve elsewhere) reaches 81.64% and 91.07%. With Open-book arithmetic as the fallback it reaches 78.56% and 86.89% (+3.07 and +4.18 pp for the program route, both intervals crossing zero). An Extract-first hybrid with a program fallback is worse by −26.44 and −25.84 pp, both clear of zero.

  • Two generic syntax lines help 7B modestly and are exploratory. The + syntax prompt (valid Python identifier names; build dates from month/day/year yourself) lifts 7B from 75.31% to 77.95%, +2.64 pp with a cluster interval of [-0.64, 6.29]; 32B moves +0.44 pp [-1.38, 2.35]. Across four checkpoints from three model families, the syntax lines shift the Program-Solve/arithmetic gap by −4.5 to +2.6 pp.

  • Scale does not predict which direction a checkpoint moves. Mistral-7B-Instruct-v0.3 scores 31.44% on formula-given code against 39.05% on arithmetic (−7.62 pp, permutation p < 0.001), while Phi-3.5-mini-instruct scores 57.13% against 52.98% (+4.15 pp, p = 0.010). Mistral's Extract-Solve library reaches 34.89%, above either Program-Solve arm.

  • Residual program failures are arithmetic hygiene, not clinical knowledge. After correcting the formula texts, no sampled error traced to an under-specified formula; the failures were unit conversions invented for supplied inputs and one-tier slips inside multi-band scoring tables. Blind failures were formula and variable errors — an ideal-body-weight convention in Cockcroft-Gault, potassium in a corrected anion gap, heart rate for respiratory rate.

  • The four guideline-replaced calculators produce the largest gap. On those four, Program-Solve scores 71.00 and 90.00 against 28.25 and 36.50 for Open-book arithmetic; removing them moves no headline gap outside its interval. Restricted to the 39 audit-clean calculators, the gap over the library is +32.26 [15.51, 48.21] at 7B and +47.59 [32.97, 61.79] at 32B.

  • Cases cluster strongly within calculators. Intraclass correlation for the library comparisons is 0.68 to 0.81, giving an effective sample size of 67 to 79 cases against 120 to 190 for the arithmetic comparisons.

Methodology in Plain English

The team took an existing clinical calculator benchmark — MedCalc-Bench Verified, 1,100 test cases spread evenly across 55 calculators at 20 cases each — and ran a set of arms that differ in exactly three things: whether the model is given the formula, whether it is handed the correct variables, and whether it does the arithmetic itself or calls something.

The decisive arm is Program-Solve. The model sees the patient note, the calculator's formula text and the correct variable values, and writes a short Python program. A subprocess runs that program in a deliberately restricted environment: only an allow-list of builtins, no file access and no eval or exec, imports limited to the standard math, date, time and calendar modules, static rejection of async and generator constructs, 256 MB memory and CPU limits, and a 5-second wall clock. The authors are explicit that this is restricted but is not a sandbox — there is no container or syscall filter.

Comparisons are matched by construction. Open-book arithmetic gets exactly the same formula, the same gold variables and the same whole-note access, but does the arithmetic in its head. Blind Program-Solve gets neither formula nor variables. The hand-written library gets the same notes but abstains outside its 22 calculators. A separate Formulate-Solve arm asks for a structured expression tree instead of code and abstains on all 60 date cases.

Two Qwen checkpoints (7B in bf16, 32B in 4-bit AWQ) were served through vLLM on cloud H100 GPUs across five seeds (42 to 46) over all 1,100 cases, with greedy decoding for single-call arms and temperature 0.7 for sampled ones. Mistral-7B and Phi-3.5-mini were run as secondary checkpoints. Because cases nest inside calculators, the primary uncertainty measure is a bootstrap that resamples calculators rather than cases.

Before all of this, the authors audited the benchmark's own formulas against current clinical guidelines, and separately audited the 55 supplied formula texts for completeness. Formula-reading arms were rerun after each fix, and the paper reports that in the final rerun every arm, including Open-book arithmetic, read the same fully corrected texts.

Why This Matters

This is one of the few studies that separates what execution buys you from what formula access and variable extraction buy you under matched conditions, and it does so while auditing whether the benchmark's clinical content is itself current. The headline is uncomfortable for both camps: handing arithmetic to a code executor helps some open-weight models and not others, and it is not a substitute for verified formulas or reliable variable extraction either way.

Real-world applications:

  • Bedside and EHR clinical decision support. Calculator selection, variable extraction and arithmetic are the three steps where LLMs lose accuracy on formulas like APACHE II; knowing that execution fixes only the third step tells implementers where to invest.
  • Drug dosing workflows. Cockcroft-Gault creatinine clearance persists in drug labelling even where CKD-EPI 2021 has replaced it for GFR estimation, so any automated dosing pipeline inherits the weight-input ambiguity the audit flags.
  • Formula versioning and governance. Race-free eGFR, MELD 3.0, PREVENT and the Sampson LDL equation have replaced predecessors that benchmarks may still reward; the paper argues formula provenance and version should be explicit inputs to any calculator interface, human or automated.
  • Local, privacy-preserving deployment. Local serving avoids external API calls and data transmission, which matters for clinical notes — though the paper notes the cost of local serving is not measured, and no Global South data, language or locale is evaluated.

Industry relevance: The findings bear on EHR vendors embedding LLM assistants, clinical-AI validation groups deciding whether to accept code execution as a safety mechanism, and anyone deploying open-weight models on-premise. The result that Mistral-7B and Phi-3.5-mini move in opposite directions on the same comparison means a checkpoint's behaviour cannot be predicted from scale, which makes per-model evaluation a prerequisite rather than an option. The security exposure of running model-written code — explicitly not sandboxed here — is a separate deployment concern raised but not resolved.

Future Directions

  1. What separates checkpoints? The paper states plainly that which of the two patterns a given open-weight checkpoint will show is not yet predictable from scale alone, and that the next question is what separates them.

  2. Formula provenance as a first-class input. Rather than inheriting version assumptions from a benchmark, future interfaces — human or automated — could take formula version and intended use as explicit inputs, especially given 16 of 55 calculators carry a version, use or coefficient concern.

  3. Separating recall from extraction. The design changes formula and variable access together in the blind arm, so it cannot tell whether a failure came from misremembering the formula or misreading the note; a controlled decomposition would need them varied independently.

  4. Locale, units and language. No Global South data, language or locale is evaluated; the benchmark is English with US conventions, and notes in other languages, laboratory values in mmol/L rather than mg/dL, and day/month/year dates each open a further path to the unit-conversion failures already observed. The exploratory date-format prompt lines show how much locale the current result silently assumes.

Target Audience

Clinical informatics researchers and hospital IT teams evaluating LLM-assisted calculators; machine-learning engineers working on tool use, code execution and program-aided reasoning; benchmark designers and clinical-AI validation groups who need to know that a benchmark's own formulas may be outdated; and health-policy or regulatory readers interested in what "verified" means when a calculator's underlying guideline has been superseded. Readers looking for a first result on whether "let the model write code" is a general fix for arithmetic unreliability will find a clear, carefully bounded negative-to-mixed answer.

Authors’ abstract

Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B and Qwen2.5-32B-AWQ, after auditing the benchmark's formulas against current clinical guidelines and flagging 16 of 55 with version, use or coefficient concerns. With formulas and gold variables supplied and both routes reading the whole note, handing off to the solver is not a reliable advantage at 7B (75.31% against 72.02%, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38]) but is one at 32B (90.53% against 83.47%, +7.05 [0.47, 14.60], clear of zero). The hand-written library is exact on its 440 supported cases but abstains elsewhere (40.0% overall). Adding an executor thus helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.

Read the original paper