Skip to content
AI.info

Research

Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study

Overview Research area: Clinical natural language processing — automated classification of cancer diagnoses from electronic health records (EHRs) using large language models and biomedical transformer

Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
arXiv
2510.12813
Published
2025-10-08
Authors
Soheil Hashtarkhani, Rezaur Rashid, Christopher L Brett, Lokesh Chinthala, Fekede Asefa Kumsa, Janet A Zink, Robert L Davis, David L Schwartz, Arash Shaban-Nejad

AI summary

Overview

Research area: Clinical natural language processing — automated classification of cancer diagnoses from electronic health records (EHRs) using large language models and biomedical transformers.

Technical level: Intermediate. The work assumes familiarity with classification metrics (weighted macro F1-score, accuracy), ICD coding, and the distinction between general-purpose LLMs and domain-specific models such as BioBERT.

Scope: A comparative evaluation of four general large language models (GPT-3.5, GPT-4o, Llama 3.2, Gemini 1.5) and one biomedical language model (BioBERT) on their ability to sort cancer diagnoses from both coded and free-text EHR entries into 14 predefined categories.

What This Paper Is About

Electronic health records store cancer diagnoses in inconsistent ways — sometimes as standardized ICD codes, sometimes as unstructured free-text notes — which makes it hard to build reliable predictive models on top of them. Automating the sorting of those diagnoses into meaningful categories with AI-driven NLP is promising, but it is unclear which models actually perform well and whether they are dependable enough for clinical work. This study compares four large language models against the biomedical model BioBERT on that classification task, using both coded and free-text diagnoses validated by oncology experts.

Key Contributions

  1. Head-to-head comparison across model families. The study evaluates four large language models (GPT-3.5, GPT-4o, Llama 3.2, Gemini 1.5) alongside the domain-specific BioBERT on a single cancer diagnosis categorization task, rather than testing one model in isolation.

  2. Separation of structured versus unstructured input. Performance is reported separately for ICD code descriptions and free-text entries, showing that the best model depends on which format the diagnosis comes in.

  3. Expert-validated classification. Two oncology experts validated the model classifications, anchoring the evaluation in clinical judgment rather than automated labels alone.

  4. Documented error patterns. The paper catalogues recurring misclassification types, notably confusion between metastasis and central nervous system tumors and errors tied to ambiguous or overlapping clinical terminology.

Main Findings

  • BioBERT led on structured ICD codes. It achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o on ICD code accuracy (90.8).
  • GPT-4o led on free-text diagnoses. For unstructured entries, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and had marginally higher accuracy (81.9 vs 81.6).
  • Smaller and older general models lagged. GPT-3.5, Gemini, and Llama showed lower overall performance on both the ICD and free-text formats.
  • Both formats are harder in free text. Even the best free-text scores (weighted macro F1 of 71.8) sit well below the best ICD-code score (84.2), indicating that unstructured diagnoses remain the tougher problem.
  • Errors cluster around clinically overlapping concepts. Common mistakes involved confusing metastasis with central nervous system tumors and mishandling ambiguous or overlapping terminology.
  • Current performance is usable for administrative and research purposes, but the authors state that reliable clinical deployment will require standardized documentation and strong human oversight for high-stakes decisions.

Methodology in Plain English

The researchers assembled a set of cancer diagnosis entries drawn from the records of patients with cancer — 762 unique diagnoses in total, split into 326 ICD code descriptions (the standardized coding format) and 436 free-text entries (how clinicians actually wrote things down), from 3,456 records overall. They defined 14 categories that each diagnosis needed to be sorted into. Then they asked five models — GPT-3.5, GPT-4o, Llama 3.2, Gemini 1.5, and BioBERT — to make those assignments. Two oncology experts checked the resulting classifications, and the team scored each model on how well its categorizations matched, reporting separate results for the coded entries and the free-text entries. They also reviewed where the models went wrong to identify recurring patterns of confusion. The abstract does not describe the prompting strategy, the exact size or composition of the evaluation split, or how the 14 categories were defined.

Why This Matters

Impact on research: The results suggest there is no single best model for EHR diagnosis categorization — a general-purpose LLM was strongest on messy free text while a smaller biomedical model was strongest on standardized codes. That finding pushes back on the assumption that scaling up general models uniformly improves clinical NLP, and it gives researchers a concrete benchmark for where domain-specific and general models each add value.

Real-world applications:

  • Cancer registry and cohort building. Automated sorting of diagnoses into categories could speed up assembling patient cohorts for research and population health tracking.
  • Administrative coding and billing workflows. The accuracy level reported for ICD code classification is described as sufficient for administrative use, which points to back-office automation.
  • Clinical documentation cleanup. Identifying where free-text entries map ambiguously to categories could flag documentation gaps for clinicians to fix.
  • Human-in-the-loop triage support. Rather than autonomous decisions, the models could pre-sort diagnoses for expert review, with oversight concentrated on the error-prone categories the study identified.

Industry relevance: Health systems, EHR vendors, and clinical AI developers evaluating whether to deploy language models for chart abstraction gain direct evidence about which model class suits which data format, and about the documentation standards and oversight needed before results can be trusted in care decisions.

Future Directions

  • Closing the free-text performance gap. The drop from roughly 84 to roughly 72 weighted macro F1 between coded and free-text input raises the question of whether better prompting, fine-tuning, or retrieval augmentation can narrow it.
  • Testing newer and larger models. The evaluated models include GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5; whether more recent releases change the ranking is an open question, as is whether models beyond this set would behave differently.
  • Addressing the specific confusion categories. The metastasis versus central nervous system tumor errors and ambiguous-terminology failures suggest targeted work on category boundaries and terminology standardization, which the abstract notes as a needed direction.
  • Establishing validation and oversight protocols. The authors call for standardized documentation practices and robust human oversight; how those safeguards should be designed and measured for high-stakes clinical use is not settled here.

Target Audience

Clinical informatics researchers and health NLP practitioners comparing language models for EHR tasks; oncologists and cancer registrars interested in what automation can and cannot currently do; health system and EHR technology teams assessing models for documentation, coding, or cohort-building pipelines; and methodologists studying how general-purpose LLMs compare with domain-specific models such as BioBERT on clinical text.

Authors’ abstract

Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise for automating diagnosis classification, their comparative performance and clinical reliability require systematic evaluation. The aim of this study is to evaluate the performance of 4 large language models (GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5) and BioBERT in classifying cancer diagnoses from structured and unstructured electronic health records data. We analyzed 762 unique diagnoses (326 International Classification of Diseases (ICD) code descriptions, 436free-text entries) from 3456 records of patients with cancer. Models were tested on their ability to categorize diagnoses into 14predefined categories. Two oncology experts validated classifications. BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in ICD code accuracy (90.8). For free-text diagnoses, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and achieved slightly higher accuracy (81.9 vs 81.6). GPT-3.5, Gemini, and Llama showed lower overall performance on both formats. Common misclassification patterns included confusion between metastasis and central nervous system tumors, as well as errors involving ambiguous or overlapping clinical terminology. Although current performance levels appear sufficient for administrative and research use, reliable clinical applications will require standardized documentation practices alongside robust human oversight for high-stakes decision-making.

Read the original paper