Skip to content
AI.info

Research

ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography

Overview Research area: Natural language processing applied to clinical cardiology, specifically domain adaptation of open-weight large language models for electrocardiography. Technical level: Interm

ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
arXiv
2510.18339
Published
2025-10-21
Authors
Lara Ahrens, Wilhelm Haverkamp, Nils Strodthoff

AI summary

Overview

Research area: Natural language processing applied to clinical cardiology, specifically domain adaptation of open-weight large language models for electrocardiography. Technical level: Intermediate — the paper assumes familiarity with finetuning, retrieval-augmented generation, and standard NLP evaluation metrics, but its framing is comparative rather than mathematically dense. Scope (one sentence): The paper finetunes Llama 3.1 8B and 70B models on ECG literature, builds matched RAG systems, and compares them against each other, their base models, and Claude Sonnet 3.7 using four distinct evaluation layers.

What This Paper Is About

Clinicians face an ever-growing volume of medical literature, and LLM question-answering systems could help them keep up — but the most capable models are closed-source, which prevents local hosting and creates privacy risks for patient data. The authors ask whether smaller, openly available models can be adapted to a narrow medical specialty (electrocardiography) well enough to compete with a leading proprietary model, and how you would even know whether they succeeded. Their central concern is that most prior domain-specialization studies rely on only one or two evaluation methods, which may hide important failure modes.

Key Contributions

  1. A direct head-to-head comparison of two domain-specialization routes — supervised finetuning and retrieval-augmented generation — applied to the same Llama 3.1 8B and 70B open-weight instruct models using the same underlying literature corpus.
  2. A multi-layered evaluation framework with four qualitatively different assessment modes (multiple-choice questions, automatic text similarity metrics, LLM-as-a-judge, and human expert evaluation), applied across all models including base versions and Claude Sonnet 3.7.
  3. Systematic RAG configuration tuning (splitting strategy, chunk size, embedding model, top-k retrieval, and reranking), reported in a dedicated appendix table, before fixing the final pipeline.
  4. An empirical bootstrapping procedure (n = 1000 iterations) that converts all evaluation modes into a unified, statistically robust ranking with explicit ties.

Main Findings

  • Finetuned 70B leads multiple-choice: Llama 3.1 70B + FT scored 90.2% on the "special" subset (1,219 samples), 92.0% on "full" (27,774 samples), and 88.2% on "checked" (534 samples), ranking first on all three. Claude Sonnet 3.7 scored 82.3%, 88.0%, and 81.7% respectively.
  • Domain-adapted models dominate multiple-choice overall: Finetuned and RAG models consistently outperformed general-purpose models, with the two domain-specialized 70B models performing best and the ranking staying consistent across all three subsets.
  • Finetuning wins on text similarity: Llama 3.1 70B + FT ranked first across ROUGE-1 (0.4270 F1), ROUGE-2 (0.2449), ROUGE-L (0.3764), BLEU (0.1289), and BERTScore (0.3904). Llama 3.1 8B + FT ranked second across all five.
  • RAG performs worst on text similarity: Llama 3.1 70B + RAG ranked last in every metric (ROUGE-1 0.2557, ROUGE-2 0.1490, ROUGE-L 0.2183, BLEU 0.0622, BERTScore 0.0181), scoring below even the base models.
  • LLM-as-a-judge favors Claude: Across 417 questions, Claude achieved the best results, closely followed by the Llama 3.1 70B models. The finetuned 8B model answered approximately 50 more questions correctly than its base model, and improvements in finetuned models over their base counterparts were statistically significant.
  • Judge reliability is asymmetric: In a human re-check, 80% of LLM-as-a-judge evaluations agreed with the human expert. All answers judged correct by the LLM were confirmed correct, but nearly 50% of answers judged incorrect by the LLM were considered correct by the human expert.
  • Human factual questions favor Claude and RAG: On ten factual questions, Claude Sonnet 3.7 and Llama 3.1 8B with RAG both answered all questions correctly. Llama 3.1 70B with RAG had one incomplete answer. No model beat those two.
  • Human complex questions favor base models and RAG: On 40 semantically complex questions, finetuned versions slightly underperformed their base models, with more completely incorrect answers. Llama 3.1 8B with RAG outperformed its base model and matched Llama 3.1 70B.
  • Complex questions are measurably harder: Average Flesch reading ease was 21.9 for the human-provided questions versus 30.3 for the training data.
  • Overall ranking: Finetuned Llama 3.1 70B had the best median rank across all categories (1.5), followed by Claude Sonnet 3.7 (2.5), then Llama 3.1 70B + RAG (3) and Llama 3.1 70B (3). Larger 70B models consistently outperformed 8B models.
  • Evaluation methods disagree: The authors emphasize substantial performance heterogeneity across evaluation methodologies, noting that no single method covers model capabilities comprehensively.

Methodology in Plain English

The authors started from ECG literature in PDF form, converted it to markdown with MinerU, and cleaned out author lists, tables of contents, references, and acknowledgements. They split files by chapter (maximum 10 chapters and 50,000 tokens estimated with TikToken) and prompted Llama 3.3 70B to generate question-answer pairs from the context, verifying context alignment with AlignScore. A modified version of the same prompt produced multiple-choice questions instead.

For finetuning, they used QLoRA via the Hugging Face Trainer on Llama 3.1 70B and 8B, with an 80/10/10 train/validation/test split per file, AdamW (paged-32), a cosine learning-rate scheduler, batch size eight, and cross-entropy loss computed only on answers. LoRA parameters were applied to attention, feedforward, and output layers, giving 3.7% trainable parameters, with r = 256 and alpha = 128. Training stopped at two epochs because validation loss rose while multiple-choice accuracy stabilized.

For RAG, they tested several configurations and settled on recursive splitting into 1,024-token chunks with 100-token overlap, PubMedBERT embeddings, top-20 retrieval, and reranking down to the top-5 chunks. They also compared multilingual-e5-large-instruct as an alternative embedding model and Chroma as the vector database.

Evaluation used four layers. Multiple-choice tests covered the "full", "special", and "checked" subsets. Automatic metrics covered BLEU, ROUGE-1, ROUGE-2, ROUGE-L, and BERTScore. LLM-as-a-judge used Deepseek R1 as the evaluator (the judge itself was not evaluated), with the original source context supplied to reduce hallucination. Human evaluation used an experienced cardiologist who labeled answers as incorrect, partially incorrect, correct but incomplete, or correct across ten factual and 40 complex questions. Rankings were derived by bootstrapping score differences 1,000 times and treating a 95% confidence interval that excludes zero as a significant difference; human categorical answers were converted to 1, 0.75, 0.25, and 0 for the four labels.

Why This Matters

The paper's central methodological message is that a domain-adaptation claim is only as strong as the evaluation behind it — and here four reasonable evaluation methods produce meaningfully different rankings. Finetuning looks dominant on in-distribution tests but loses its edge on syntactically and semantically unfamiliar questions, while RAG looks weak on lexical metrics yet matches or beats finetuning in human evaluation. That divergence is a warning for anyone drawing conclusions from a single benchmark.

Real-world applications the authors describe:

  • Real-time emergency interpretation for distinguishing ST-elevation myocardial infarction (STEMI) mimics from true infarction.
  • Integration into continuous monitoring for intelligent alarm management and arrhythmic event prediction.
  • Rapid processing of large-scale ECG databases to identify novel biomarkers.
  • Point-of-care, evidence-based guidance supporting arrhythmia identification, ST-segment analysis, and QT assessment.

Industry relevance: The viability of privacy-preserving, locally deployable clinical tools matters commercially as well as clinically. A locally hosted 8B model augmented with RAG reaching the performance of much larger systems, as reported here, changes the cost and compliance calculus for hospital systems that cannot send patient data to external APIs. The authors cite ChatGPT-4o achieving 75.9% accuracy on cardiology board questions as context for how general-purpose tools currently perform in this space. They also caution that deployment requires real-world testing, safety evaluations, workflow integration, and regulatory approval, and that at this stage the models are better suited to helping clinicians find information than to making autonomous decisions.

Future Directions

  • Combine RAG with finetuning. The authors flag this as the most obvious untested combination, and note that it is unclear whether finetuned models can effectively use retrieved contextual information or would need a different finetuning strategy to do so.
  • Broaden and complicate the training data. More varied tasks (diagnoses, recommendations, complex relationships), human-generated data, and complex chat instructions could address the out-of-distribution weakness the human evaluation exposed. The authors also note that synthetic data generation via LLM prompting lacks control over diversity and complexity.
  • Try continual pretraining before supervised finetuning, which one cited study supports but another questions — the authors leave this unresolved and also suggest RLHF as a route to align models with defined clinical preferences.
  • Evaluate with pairwise preference comparisons such as Elo-style scoring, which this work deliberately did not include since the models already differed substantially in factual correctness.
  • Improve the RAG pipeline through deeper analysis of embedding models, finetuning embeddings, and curating document selection.

Target Audience

Clinical informatics and NLP researchers working on domain adaptation or medical question answering will find the evaluation-framework design and the finetuning-versus-RAG comparison most useful. Cardiologists and electrophysiologists interested in what current LLM tooling can and cannot do for ECG interpretation will benefit from the human evaluation results specifically. Machine learning engineers building locally deployable clinical systems will find the concrete finetuning hyperparameters (r = 256, alpha = 128, QLoRA, 3.7% trainable parameters) and the final RAG configuration directly actionable. Researchers designing benchmark studies more broadly should read it for the bootstrapping-based ranking methodology and the demonstration that evaluation-mode choice can flip conclusions.

Authors’ abstract

Domain-adapted open-weight large language models (LLMs) offer promising healthcare applications, from queryable knowledge bases to multimodal assistants, with the crucial advantage of local deployment for privacy preservation. However, optimal adaptation strategies, evaluation methodologies, and performance relative to general-purpose LLMs remain poorly characterized. We investigated these questions in electrocardiography, an important area of cardiovascular medicine, by finetuning open-weight models on domain-specific literature and implementing a multi-layered evaluation framework comparing finetuned models, retrieval-augmented generation (RAG), and Claude Sonnet 3.7 as a representative general-purpose model. Finetuned Llama 3.1 70B achieved superior performance on multiple-choice evaluations and automatic text metrics, ranking second to Claude 3.7 in LLM-as-a-judge assessments. Human expert evaluation favored Claude 3.7 and RAG approaches for complex queries. Finetuned models significantly outperformed their base counterparts across nearly all evaluation modes. Our findings reveal substantial performance heterogeneity across evaluation methodologies, underscoring assessment complexity. Nevertheless, domain-specific adaptation through finetuning and RAG achieves competitive performance with proprietary models, supporting the viability of privacy-preserving, locally deployable clinical solutions.

Read the original paper