Skip to content
AI.info

Research

Multi-Window Temporal Analysis for Enhanced Arrhythmia Classification: Leveraging Long-Range Dependencies in Electrocardiogram Signals

Overview Research area: Deep learning for biomedical time-series analysis, specifically automated arrhythmia classification from electrocardiogram (ECG) recordings, with a focus on atrial fibrillation

Multi-Window Temporal Analysis for Enhanced Arrhythmia Classification: Leveraging Long-Range Dependencies in Electrocardiogram Signals
arXiv
2510.17406
Published
2025-10-20
Authors
Tiezhi Wang, Wilhelm Haverkamp, Nils Strodthoff

AI summary

Overview

Research area: Deep learning for biomedical time-series analysis, specifically automated arrhythmia classification from electrocardiogram (ECG) recordings, with a focus on atrial fibrillation (AF) detection and cross-dataset generalization.

Technical level: Intermediate. The paper assumes familiarity with deep learning concepts (encoders, sequence models, AUROC) but explains the clinical motivation and architectural choices in accessible terms.

Scope: A systematic study of whether jointly analyzing multiple consecutive 30-second ECG windows—rather than isolated single windows—improves arrhythmia classification accuracy, false-positive rates, and robustness to domain shift, using a structured state-space (S4) architecture called S4ECG.

What This Paper Is About

Most automated ECG arrhythmia detectors analyze a single 5–30 second window in isolation, which limits the temporal context available for recognizing rhythm transitions, paroxysmal episodes, and gradual changes that unfold over minutes. This causes high false positive rates (reported specificity of 0.72–0.98 for AF detection with conventional 30-second windows) and poor generalization when models are tested on data from other institutions or acquisition protocols. The paper's goal is to build a model that analyzes many consecutive windows at once (up to 20 minutes of continuous ECG) and to systematically determine how much temporal context is actually needed.

Key Contributions

  1. Adaptation of an encoder-predictor paradigm to ECG. The authors extend a hierarchical architecture previously developed for sleep staging (Wang and Strodthoff, 2025a) to arrhythmia detection, using S4 layers at both the window level and the multi-window sequence level, creating the S4ECG model.
  2. Evidence that multi-window prediction beats single-window prediction. Across four datasets and both in-distribution and out-of-distribution (OOD) evaluation, jointly predicting multiple consecutive windows improved macro-averaged AUROC by 1.0–11.6 percentage points over single-window baselines.
  3. The first systematic investigation of the optimal temporal window for ECG arrhythmia detection. The authors sweep temporal contexts from 2 to 60 windows (1 to 30 minutes at a fixed 30-second window size) for LTAFDB and 10 to 60 windows (5 to 30 minutes) for Icentia11k, finding a consistent optimum in the 20–40 window range (10–20 minutes).
  4. A rigorous cross-dataset robustness evaluation. Models trained on Icentia11k or LTAFDB are tested on external PhysioNet databases (AFDB, MITDB, LTAFDB) with different patient populations, sampling rates, and acquisition protocols, showing that multi-window models degrade less under domain shift.

Main Findings

  • S4 encoders outperform CNN encoders. On Icentia11k, the single-window S4 model reached a macro-AUROC of 0.9702 versus 0.9663 for the xResNet1d50 baseline (a 0.4% improvement). On LTAFDB the gap was larger: 0.9029 for S4 versus 0.8621 for ResNet (a 4.7% improvement).

  • In-distribution multi-window gains. On Icentia11k, the best configuration used 30 input windows with a macro-AUROC of 0.9800 (+1.0% over the single-window S4 model), and AF specificity at a fixed sensitivity of 0.9 rose from 0.9033 to 0.9869. On LTAFDB, performance peaked at 10 input windows with a macro-AUROC of 0.9684 (+7.3%), and AF AUROC improved from 0.8319 to 0.9841 (+18.3%).

  • A stable 20–40 window operating range. On LTAFDB, macro-AUROC remained between 0.9612 and 0.9684 across 10 to 40 windows, with the 20-window configuration (0.9664) statistically equivalent to the best model. Different rhythm classes peaked at different points: normal rhythm (N) at 20 windows (0.9950) and SVTA at 40 windows (0.9285).

  • Strong false-positive reduction in AF detection. AF specificity at sensitivity 0.9 rose from 0.9794 (single-window S4) to 0.9983 at 30 windows on LTAFDB, and from 0.9033 to 0.9869 at 30 windows on Icentia11k. Across the abstract's summary, single-window specificity spanned 0.718–0.979 versus 0.967–0.998 for multi-window, described as a 3–10 fold reduction in false positive rates.

  • OOD generalization improved substantially. For a model trained on Icentia11k, AFDB macro-AUROC peaked at 30 windows (0.9328 versus 0.8718 single-window, +7.0%), with AF specificity rising from 0.9572 to 0.9998 and atrial flutter AUROC improving from 0.6899 to 0.8190 (+18.7%).

  • Largest OOD gains on MITDB. MITDB macro-AUROC improved from 0.8426 (single-window) to 0.9401 at 30 windows (+11.6%), with the 40-window configuration statistically equivalent (0.9382). Normal rhythm AUROC rose from 0.6983 to 0.8465 (+21.2%), and AF specificity from 0.7962 to 0.9743 at 40 windows.

  • Icentia11k-to-LTAFDB transfer. The 30-window model achieved a macro-AUROC of 0.9767 (+5.5% over single-window), and AF specificity improved from 0.718 to 0.9884 at 20 windows, reported as a 37.6% increase. AF AUROC rose from 0.9427 to 0.9748 (+3.4%).

  • Performance is non-monotonic in sequence length. Gains are rapid up to moderate lengths, then plateau, with degradation beyond 40 windows. The authors attribute the decline primarily to optimization difficulty on longer sequences (a more complex loss landscape with local minima) rather than architectural limits, citing prior sleep-staging work in which curriculum training recovered performance lost under direct long-sequence training. Secondary factors include diminishing diagnostic returns beyond the physiologically meaningful ~10–20 minute timescale.

  • Qualitative temporal coherence. On a continuous AF episode from LTAFDB, the 30-window model (trained on Icentia11k) produced predictions matching the ground-truth annotation and estimated AF burden more accurately, while the single-window model produced fragmented predictions with spurious interruptions.

  • Clinical and physiological grounding. The paper cites prior evidence that healthy heart dynamics exhibit multifractal complexity persisting for at least 700 beats (~10 minutes) and that long-range correlations in 24-hour recordings span roughly 10^2 to 10^3 beats (~1 to 20 minutes), whereas pathological dynamics deviate at those scales.

Methodology in Plain English

The researchers built a two-stage model. In the first stage, each 30-second ECG window (3,840 samples at 128 Hz) is passed through a small convolutional front-end that reduces it from 3,840 to 960 samples, then through a stack of S4 layers (model dimension 512, state dimension 64, four layers, bidirectional). A pooling step compresses each window into a single 512-dimensional token. In the second stage, the sequence of these tokens is fed into a second four-layer S4 module (model dimension 512, bidirectional) that learns relationships across windows, and a linear classification head produces a rhythm prediction for every window. For example, an input of 38,400 samples corresponds to 10 windows spanning 5 minutes.

The models were trained directly on raw ECG time series with no baseline-wander correction or filtering, so the network learns from unfiltered signals and avoids information loss from hand-crafted preprocessing. Training targets are fractional labels: the proportion of each rhythm type within each prediction window, optimized with a binary cross-entropy loss using the AdamW optimizer.

Evaluation covered four public PhysioNet databases: Icentia11k (11,000 patients, 250 Hz, ~110,000 hours) and LTAFDB (84 patients, 128 Hz, 24–25 hours per recording, ~2,000 hours total) for training and in-distribution testing, with AFDB (25 patients, 250 Hz, ~250 hours) and MITDB (47 patients, 360 Hz, 30-minute recordings, ~24 hours) used for out-of-distribution testing. The primary metric is macro-averaged AUROC; AF specificity is reported at a fixed sensitivity of 0.9, matching the operating point of FDA-cleared wearable AF detectors. Statistical significance was assessed with a patient-level paired bootstrap using 10,000 resamples and 95% confidence intervals. Two single-window baselines were compared: the xResNet1d50 CNN and the strongest S4-based single-window model from prior work.

Why This Matters

Impact on research. The paper challenges the dominant single-window paradigm in ECG deep learning and provides the first systematic sweep of temporal context lengths for arrhythmia detection. It argues for a methodological shift both in temporal modeling (single-window to multi-window) and in architecture (CNN/LSTM to structured state-space models that scale linearly with sequence length). The finding that a 20–40 window optimum in ECG mirrors moderate-sequence-length results in sleep staging suggests a more general design principle for physiological time-series analysis.

Real-world applications:

  • Remote and wearable cardiac monitoring, where false alarms are a major burden—false positives account for almost 60% of overall remote transmissions from implantable loop recorders, and false positives from consumer-grade devices contribute to increased emergency department utilization. Improved specificity at fixed sensitivity could reduce these.
  • Atrial fibrillation burden quantification, which is used for stroke risk assessment; the paper notes that even modest AF burden above 0.5% correlates with increased thromboembolic risk.
  • Detection of short paroxysmal episodes (seconds to minutes) that conventional monitoring misses, relevant for post-ablation monitoring and for patients with cryptogenic cerebrovascular events.
  • Analysis of rhythm transitions and mode switching (for example, AF to atrial flutter), which can inform catheter ablation strategy and reveal mechanistic information about arrhythmia maintenance and termination.

Industry relevance. The robustness results speak directly to regulatory expectations: the paper cites the FDA's Software as a Medical Device guideline and its action plan on AI/ML-based SaMD, both of which emphasize performance stability under domain shift. Because S4 models scale linearly with sequence length, the approach is computationally attractive for edge devices, though the authors state that thorough assessment in wearable and mobile health edge-computing environments is still needed.

Future Directions

  1. Self-supervised and semi-supervised learning. The study is limited to supervised training requiring extensive expert-labeled data. The multi-window design is described as naturally compatible with self-supervised objectives that exploit inter-window temporal relations, which could harness the large unlabeled ECG datasets increasingly common in practice.
  2. Real-time and edge deployment. All evaluation was retrospective; validation in real-time clinical monitoring systems and on wearable or mobile edge hardware remains to be established.
  3. Adaptive sequence lengths. The current model processes fixed-length sequences, which may not optimally capture variability in arrhythmic episode duration. Mechanisms that adjust context based on rhythm stability are proposed as a potential improvement.
  4. Resolving the long-sequence optimization bottleneck. Since the authors attribute performance degradation beyond 40 windows to optimization rather than architecture, curriculum training approaches—progressively extending sequence length during training—remain an open avenue for extending useful temporal context further.

Target Audience

This paper is most valuable to machine learning researchers working on biomedical time series and sequence modeling, particularly those interested in state-space models as alternatives to CNNs and LSTMs. It also serves clinical and regulatory audiences evaluating

Authors’ abstract

Objective. Arrhythmia classification from electrocardiograms (ECGs) suffers from high false positive rates and limited cross-dataset generalization, particularly for atrial fibrillation (AF) detection where specificity ranges from 0.72 to 0.98 using conventional 30-s analysis windows. While most deep learning approaches analyze isolated 30-s ECG windows, many arrhythmias, including AF and atrial flutter, exhibit diagnostic features that emerge over extended time scales. Approach. We introduce S4ECG, a deep learning architecture based on structured state-space models (S4), designed to capture long-range temporal dependencies by jointly analyzing multiple consecutive ECG windows spanning up to 20 min. We evaluate S4ECG on four publicly available databases for multi-class arrhythmia classification and perform systematic cross-dataset evaluations to assess out-of-distribution robustness. Results. Multi-window analysis consistently outperforms single-window approaches across all datasets, improving macro-averaged AUROC by 1.0-11.6 percentage points. For AF, specificity increases from 0.718-0.979 to 0.967-0.998 at a fixed sensitivity threshold, yielding a 3-10-fold reduction in false positive rates. Significance. Compared with convolutional neural network baselines, the S4 architecture shows superior performance, and multi-window training substantially reduces cross-dataset degradation. Optimal diagnostic windows are 10-20 min, beyond which performance plateaus or degrades. These findings demonstrate that structured incorporation of extended temporal context enhances both arrhythmia classification accuracy and cross-dataset robustness. The identified optimal temporal windows provide practical guidance for ECG monitoring system design and may reflect underlying physiological timescales of arrhythmogenic dynamics.

Read the original paper