Skip to content
AI.info

Research

A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients

Overview Research area: Machine learning for clinical electrocardiography — specifically automated detection of atrial fibrillation (AF) from ECG signals recorded in intensive care units (ICUs). Techn

A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients
arXiv
2512.18031
Published
2025-12-19
Authors
Sarah Nassar, Nooshin Maghsoodi, Sophia Mannina, Shamel Addas, Stephanie Sibley, Gabor Fichtinger, David Pichora, David Maslove, Purang Abolmaesumi, Parvin Mousavi

AI summary

Overview

Research area: Machine learning for clinical electrocardiography — specifically automated detection of atrial fibrillation (AF) from ECG signals recorded in intensive care units (ICUs).

Technical level: Intermediate. The paper assumes familiarity with standard ML concepts (classifiers, cross-validation, F1 score) and basic ECG terminology, but the benchmarking framework itself is straightforward to follow.

Scope: The paper benchmarks three families of AI models — feature-based classifiers, deep learning models, and ECG foundation models — across four training configurations on a newly published ICU ECG dataset and the 2021 PhysioNet/CinC Challenge dataset, and releases the labelled ICU dataset publicly.

What This Paper Is About

Atrial fibrillation is the most common cardiac arrhythmia in ICU patients, with a prevalence reported in the paper of up to 15-20% or more, and ICU patients are continuously monitored, making automated detection attractive. Despite a large literature on AI-based arrhythmia detection, almost all of it uses deep learning, and there has been no direct, comprehensive comparison of classical machine learning, deep learning, and the newer ECG foundation models — especially on ICU data, which is noisy and uses a reduced three-to-five-lead telemetry setup rather than a standard 12-lead snapshot. This paper fills that gap by benchmarking many models under multiple training regimes and by publishing a labelled ICU dataset to support further work.

Key Contributions

  1. A comprehensive benchmark of AF detection across three data-driven AI approaches (feature-based classifiers, deep learning, and foundation models) and four training configurations, using both the large public 2021 PhysioNet/CinC Challenge dataset and the authors' institutional ICU dataset.
  2. Public release of the institutional ICU dataset — almost 600 labelled 10 s four-lead ECGs from unique patients from Kingston, Ontario, Canada — available at https://physionet.org/content/kingston-icu-af/.
  3. A demonstration that fine-tuning an ECG foundation model improved AF detection performance on the ICU dataset compared to previous research on that data.
  4. External validation of the observed trends on a non-ICU dataset (2021 PhysioNet), showing the ranking of approaches is broadly consistent across domains.

Main Findings

  • Overall ranking of approaches: Across both datasets, ECG foundation models generally performed best, followed by deep learning, then feature-based classifiers. However, the gap between DL and feature-based classifiers was minimal and highly dependent on the specific model chosen. Feature-based classifiers achieved a very slightly higher average performance on the ICU test set, while DL achieved a higher overall maximum performance.

  • Top ICU test set performance: The models achieving the top F1 score of 0.88 on the ICU test set were InceptionV3 with recurrence plots (training config #3, also best precision of 0.98) and a fine-tuned ECGFounder (training config #4, also best precision of 0.98).

  • Zero-shot foundation models (config #1, n=0): The best F1 score of 0.86 was achieved by ECG-FM v1 with no additional training. ECG-FM v1 outperformed ECG-FM v2 across all metrics except recall. The 12-lead and single-lead ECGFounder variants had very similar F1 scores.

  • Training on ICU data only (config #2, n=298): The best F1 score of 0.85 and best precision of 0.89 came from a fine-tuned ECGFounder. All feature-based models and both fine-tuned ECG foundation models had higher F1 and precision than all 1-D and 2-D DL models trained from scratch, which frequently produced all-negative or all-positive predictions (leading to zero or undefined metrics) because they had too many parameters for the small training set.

  • Training on PhysioNet data (config #3, n=70,363): With tens of thousands of samples, many 1-D and 2-D DL models surpassed the feature-based models, and fine-tuned ECG foundation models also outperformed them. The best recall of 0.84 was achieved by ResNet-18 with ECG images or recurrence plots and by a fine-tuned ECG-FM.

  • Input length matters for 1-D CNNs: The Goodfellow CNN's F1 rose from 0.77 to 0.85 and precision from 0.73 to 0.89 when input size increased from 2.5 s to 10 s. The Stanford CNN's F1 rose from 0.82 to 0.85 and precision from 0.82 to 0.91 when input size increased from 1.1 s to 9.6 s.

  • Transfer learning (config #4, n=70,363+298): Transfer learning did not prove advantageous compared to config #3 except for ResNet-18 with spectrograms or scaleograms and ECGFounder. The best recall of 0.88 came from the 1-D Stanford CNN with 9.6 s inputs.

  • Signal-to-image transforms mostly underperformed: With the exception of Inception v3 with recurrence plots, the 2-D CNNs using spectrograms, scaleograms, and recurrence plots generally did not reach high performance.

  • Feature-based vs DL nuance: In the paper's own comparison on the ICU test set, the best feature-based model reached F1=0.84 versus F1=0.88 for the best DL model; on the 2021 PhysioNet data the comparison was F1=0.89 vs F1=0.91. The authors note this gap is much smaller than the F1=0.700 vs F1=0.869 gap reported by one prior study, suggesting feature-based performance depends heavily on the quality of the hand-crafted features.

  • Improvement over prior Kingston ICU results: Previous work on the same data reported average F1 up to 0.78 with 2.5 s windows and F1 up to 0.67 or 0.714 with 10 s inputs using a 1-D CNN similar to the Goodfellow CNN. In this work, the Goodfellow CNN reached a maximum F1 of 0.77 with 2.5 s windows and 0.85 with 10 s inputs.

  • Confidence intervals: The 23 models with the best F1 scores were evaluated by bootstrapping each applicable test set 10,000 times, with 95% confidence intervals reported in the supplementary materials.

  • Dataset scarcity for ICU ECGs: The authors state that the MIMIC-IV-ECG database (about 800,000 ECGs from nearly 160,000 ICU patients with more than 600,000 cardiologist reports) has not published de-identified free-text cardiologist reports, so to their knowledge no other open-source dataset with labelled ECGs from ICU patients exists.

  • Human comparison result: The paper cites an average cardiologist F1 score of 0.677 from the Stanford CNN paper, evaluated on 328 30 s ECG recordings from distinct patients; the supplied text is truncated before this discussion concludes.

Methodology in Plain English

The researchers assembled two datasets. The first is their own ICU dataset from a tertiary hospital in Kingston, Ontario, Canada, containing archival data from 2015 to 2020 from bedside monitors covering 1,043 de-identified patients, with telemetry leads I, II, III, and a V1 equivalent sampled at 240 Hz. From a subset of patients, a randomly selected 10 s ECG was annotated by two critical care physicians, with a third consulted in case of ties, yielding 613 relevant ECGs where 513 were labelled sinus rhythm and 100 were labelled AF (combined with atrial flutter). After dropping segments that could not yield features, 596 ECGs remained with 17% AF prevalence. The second dataset is the training set of the 2021 PhysioNet/CinC Challenge, covering 70,363 ECGs with 18% overall AF prevalence, drawn from four source sites (Chapman/Shaoxing & Ningbo, CPSC_2018 & CPSC_2018_Extra, Georgia, and PTB & PTB-XL). SNOMED CT codes were mapped into a sinus rhythm class and an AF class, recordings with both labels or with co-occurring unused labels were removed, flat or missing leads were excluded, signals were bandpass filtered (5-30 Hz Butterworth, second order), and the public data was downsampled to 240 Hz to match the ICU data.

For the feature-based approach, R peaks were detected with the SleepECG package using a modified adaptive-threshold beat detector, and NeuroKit2 was used to compute signal quality, heart rate, heart rate variability features (time, non-linear, and fragmentation measures), entropy, and a P peak missingness feature — 27 features per lead across four leads, giving 108 columns per recording. These were fed to KNN, MLP, SVM, logistic regression, and four tree ensembles with 1,000 estimators each (random forest, LightGBM, XGBoost, CatBoost), plus zero-shot TabPFN v2, after min-max scaling.

For deep learning, the authors tested 1-D CNNs on raw signals (a Goodfellow-style CNN and the Stanford CNN), 2-D CNNs on plotted ECG images (ResNet-18 and Inception v3, image size 319 × 999), and 2-D CNNs on signal-to-image transformations (spectrograms from the short-time Fourier transform, scaleograms from the continuous wavelet transform, and recurrence plots), with the four lead images combined as channels by modifying the first convolutional layer.

For foundation models, they used two open-weight pretrained models that take raw signals: the transformer-based ECG-FM (pretrained on 1.4 million 10 s 12-lead ECGs from MIMIC-IV-ECG and the 2021 PhysioNet Challenge) and the CNN-based ECGFounder (pretrained on over 7.5 million 10 s 12-lead ECGs from the Harvard-Emory ECG Database). Since the ICU data has only four leads, the input layer was modified to accept four leads while keeping pretrained weights for the relevant leads.

The ICU dataset was split in a stratified way into a training set (n=298) and a test set (n=298). Models were compared under four configurations: (1) zero-shot inference with no training, (2) training on the ICU data, (3) training on the PhysioNet data with leave-site-out cross-validation to avoid patient leakage, and (4) transfer learning — training on PhysioNet then fine-tuning on the ICU training set. Learning rates were 0.05 for the 1-D and 2-D CNNs and 0.0001 for foundation model fine-tuning, batch sizes were 32, 64, or 128, training used binary cross-entropy loss and the Adam optimizer, early stopping was based on the ICU training set for up to 200 epochs, and the classification threshold was 0.5. Recall, precision, and F1 were the comparison metrics, with F1 chosen as the primary metric to balance the need to catch AF events against the need to avoid false alarms.

Why This Matters

Impact on research. The paper provides the first broad, apples-to-apples comparison of classical machine learning, deep learning, and ECG foundation models for AF detection, on both an ICU dataset and a large public dataset. It challenges the implicit assumption in the literature that deep learning is automatically the best choice, and it releases a labelled ICU dataset at a time when the authors report no other open-source labelled ICU ECG dataset is available. It also establishes performance baselines against which future work can be measured.

Real-world applications:

  • Continuous automated cardiac rhythm monitoring at the bedside in ICUs, where patients are already connected to telemetry monitors.
  • Alarm-triage support, since the paper emphasizes that high precision is needed to avoid exacerbating alarm fatigue in care providers.
  • Resource-constrained deployment scenarios, where lighter feature-based classifiers may be viable alternatives to large deep networks.
  • Supporting clinicians with a second opinion on ECG interpretation, given that the paper notes ICU ECGs are continuous and interpretation is time-intensive and requires experience.

Industry relevance. Medical device and patient-monitoring companies looking to add automated arrhythmia detection to bedside telemetry systems; healthcare AI vendors evaluating whether to invest in foundation model fine-tuning versus lighter-weight feature-based pipelines; and health systems considering compute costs, inference speed, and false-alarm rates as selection criteria — factors the authors explicitly say can decide model choice once top F1 scores converge.

Future Directions

  • AF forecasting. The authors identify future prediction of AF episodes (rather than only flagging current AF) as the latest research frontier, and propose that high-performing detection models could generate weak labels across entire patient stays from long-term recordings, since it is not feasible for an expert annotator to process hours or days of data.
  • Larger ICU datasets. The paper highlights the scarcity of AF studies in the ICU with datasets spanning enough patients, and notes the absence of an alternative open-source labelled ICU ECG dataset, leaving room for larger multicentre ICU collections.
  • Deployment-oriented model selection. With top F1 scores closely clustered, the authors suggest further work on choosing models based on precision, inference speed, and computational requirements for real-time implementation with minimal false alarms.
  • Broader and deeper comparative benchmarking. The authors note that their results differ from some earlier comparisons and attribute this partly to the variety and number of approaches tested, suggesting larger-scale comparison experiments may yield more consistent general trends across datasets.

Target Audience

This paper is most useful to machine learning researchers working on clinical time-series and ECG analysis; biomedical engineers and clinical informatics teams building ICU monitoring systems; critical care physicians and cardiologists interested in what current AI can and cannot do for AF detection in their patients; and data scientists at medical device or health technology companies evaluating model families for real-time arrhythmia detection. It is also valuable to researchers specifically interested in ECG foundation models, since it includes ECGFounder and ECG-FM in a direct comparison with conventional approaches.

Authors’ abstract

Objective: Atrial fibrillation (AF) is the most common cardiac arrhythmia experienced by intensive care unit (ICU) patients and can cause adverse health effects. In this study, we publish a labelled ICU dataset and benchmarks for AF detection. Methods: We compared machine learning models across three data-driven artificial intelligence (AI) approaches: feature-based classifiers, deep learning (DL), and ECG foundation models (FMs). This comparison addresses a critical gap in the literature and aims to pinpoint which AI approach is best for accurate AF detection. Electrocardiograms (ECGs) from a Canadian ICU and the 2021 PhysioNet/Computing in Cardiology Challenge were used to conduct the experiments. Multiple training configurations were tested, ranging from zero-shot inference to transfer learning. Results: On average and across both datasets, ECG FMs performed best, followed by DL, then feature-based classifiers. The model that achieved the top F1 score on our ICU test set was ECG-FM through a transfer learning strategy (F1=0.89). Conclusion: This study demonstrates promising potential for using AI to build an automatic patient monitoring system. Significance: By publishing our labelled ICU dataset (LinkToBeAdded) and performance benchmarks, this work enables the research community to continue advancing the state-of-the-art in AF detection in the ICU. https://physionet.org/content/kingston-icu-af/

Read the original paper