Skip to content
AI.info

Research

CLEF: Clinically-Guided Contrastive Learning for Electrocardiogram Foundation Models

Overview Research area: Machine learning for healthcare — self-supervised pretraining of foundation models for electrocardiogram (ECG) signals, with a focus on single-lead ECGs used in wearables and r

CLEF: Clinically-Guided Contrastive Learning for Electrocardiogram Foundation Models
arXiv
2512.02180
Published
2025-12-01
Authors
Yuxuan Shu, Peter H. Charlton, Fahim Kawsar, Jussi Hernesniemi, Mohammad Malekzadeh

AI summary

Overview

  • Research area: Machine learning for healthcare — self-supervised pretraining of foundation models for electrocardiogram (ECG) signals, with a focus on single-lead ECGs used in wearables and remote monitoring.
  • Technical level: Intermediate. The paper assumes familiarity with contrastive learning and deep learning for time series, but the core idea (using a clinical risk score to shape representation learning) is explained in accessible terms.
  • Scope: The paper introduces CLEF, a clinically-guided contrastive pretraining method that uses the SCORE2 cardiovascular risk score derived from routine patient metadata to weight negative pairs, and evaluates it on 18 clinical classification and regression tasks across 7 held-out datasets.

What This Paper Is About

Existing ECG foundation models are pretrained with self-supervised methods that treat all negative pairs equally, ignoring clinical knowledge about how different two patients actually are. The authors propose using an established, clinically validated risk score (SCORE2) computed from routinely collected metadata — age, gender, smoking status, systolic blood pressure, diabetes status, total cholesterol, and HDL cholesterol — to control how strongly ECG embeddings are pushed apart during contrastive pretraining, with an explicit mechanism for handling missing metadata. The goal is to build single-lead ECG models that learn clinically meaningful representations without requiring per-sample ECG annotations, and that transfer to wearable and emergency-care settings.

Key Contributions

  1. Clinically-guided contrastive learning for ECG representation learning. The method uses a clinically validated risk score to define soft, continuous similarity relationships between unlabeled ECG samples, rather than the binary similar/dissimilar notion used in classic contrastive learning. The authors state this lets the model learn richer, more nuanced representations, and includes a mechanism to handle missing metadata.
  2. Extensive empirical analysis and benchmarking for ECG foundation models. The evaluation covers (i) downstream performance on lead I/II ECG classification and regression, (ii) comparison against 12-lead ECG approaches, (iii) robustness across different model architectures and pretraining methods, (iv) ablation studies on loss components and handling missing metadata, and (v) linear probing to assess representation quality.
  3. Effectiveness across diverse downstream tasks. Evaluated on a wide range of applications, CLEF is reported to consistently outperform strong baselines, with an average improvement of 3.1% in classification AUROC and a reduction of 2.9% in regression MAE.
  4. Model release. Code and pretrained CLEF models are made available at github.com/Nokia-Bell-Labs/ecg-foundation-model.

Main Findings

  • Pretraining scale and setting: CLEF is pretrained on 12-lead ECGs from 161,352 patients in MIMIC-IV-ECG, using only routinely collected metadata. Three single-lead model scales are built with ResNeXt1D: CLEF-S (448K parameters), CLEF-M (30.7M parameters), and CLEF-L (296M parameters).
  • Headline comparisons: When pretrained on 12-lead data and tested on lead-I data, the medium CLEF achieves average AUROC improvements of at least 2.6% in classification and average MAE reductions of at least 3.2% in regression. Against existing self-supervised learning algorithms, CLEF improves average AUROC by at least 1.8%.
  • Lead I and lead II fine-tuning: The best CLEF variant for each task improves average AUROC by 3.1% on lead I and 4.8% on lead II over the strongest baseline for that task. Individually, CLEF-S, -M, and -L improve by 1.0%, 2.6%, and 1.3% over the best baseline on lead I (1.6% on average), and by 2.8%, 4.0%, and 4.2% on lead II (3.7% on average). KED was the best-performing baseline in all but one task, and CLEF-M shows statistically significant superiority over KED (p = 0.010, paired t-test).
  • Margin over the strongest ECG baseline: CLEF outperforms KED by at least 1.0% on lead I and at least 2.8% on lead II across all 3 variants. The averaged confidence interval across all CLEF variants and tasks is ±0.01.
  • Strength on pattern-identification tasks: On the PTB-XL form and rhythm tasks and the Chapman arrhythmia task, the best CLEF variant improves AUROC by 5.9% (lead I) and 8.5% (lead II) on average over the strongest baselines, which the authors attribute to capturing both the morphology of individual beats and long-range rhythm dependencies.
  • Cross-lead robustness: CLEF-S, -M, and -L differ by 0.3%, 0.0%, and 1.4% on average in AUROC across the 7 tasks when moving between leads. By contrast, ST-MEM and KED favored lead I, with lead II AUROCs averaging 4.2% and 1.5% lower, respectively.
  • Comparison with the supervised model ECGFounder: CLEF performs less well on lead I (AUROCs lower by 2.9%, 1.5%, and 2.7% for CLEF-S, -M, and -L) but better on lead II (AUROC increases of 1.3%, 2.5%, and 2.6%). The paper attributes this to ECGFounder dropping 3.8% on lead II relative to lead I, the lead on which the single-lead version of ECGFounder was trained. CLEF-S also outperforms ST-MEM on 5 of 7 tasks even though ST-MEM was pretrained and evaluated on 12-lead ECGs.
  • Single-lead, out-of-clinic classification: On MUSIC, Icentia11K, and MC-MED, the best CLEF variant per task improves AUROC by 2.5% on average over the strongest baseline. CLEF-S, -M, and -L achieve average improvements of 2.8%, 6.7%, and 2.5% over the best baseline (KED). On the Icentia11K Beat and Rhythm tasks, the best CLEF variant per task improves AUROC by 1.9% over the best baselines, with CLEF-S, -M, and -L improving by 1.5%, 2.5%, and 3.7% on average over KED.
  • One case where a general time series model won: Moment outperformed CLEF on the MUSIC sudden cardiac death (SCD) prediction task, the only case in which a general time series model surpassed an ECG-specific model. The authors suggest SCD risk depends on long-range, subtle temporal patterns, making it more like forecasting than short-term classification.
  • Regression tasks: The best CLEF variant per task achieves an average MAE reduction of 2.9% versus the strongest baseline. CLEF-S and CLEF-M outperform all baselines with average MAE reductions of 3.2% versus the best baseline (ST-MEM); CLEF-L beats 3 of 4 baselines but has an MAE 2.3% higher than ST-MEM on average. CLEF-S, -M, and -L achieve lower MAEs than the supervised ECGFounder by 5.5%, 5.6%, and 0.2% on average.
  • Linear probing: With frozen encoders, the best CLEF per task outperforms the strongest baseline for that task by +7.3%. CLEF-M and CLEF-L both beat all baselines, with average AUROC improvements of 8.5% and 9.9% over the best baseline (KED). CLEF-S does not outperform KED or Moment, indicating larger CLEF models produce better quality representations. CLEF-M outperforms Moment by 10.0% and CLEF outperforms ST-MEM by 28.1%.
  • Comparison with self-supervised algorithms: Against SimCLR, BYOL, and MoCo trained on all three backbone sizes, CLEF outperformed 100 out of 117 instances, with average AUROC improvements of 29.8%, 1.8%, and 2.3% versus MoCo, SimCLR, and BYOL respectively.
  • Applying the objective to an existing baseline: Initializing KED with its original pretrained weights and then applying the clinically-guided pretraining improved every task except LVEF classification on MIMIC-IV, with an average AUROC gain of 3.0%; the PTB-XL form task improved the most, by 11.7%.
  • Pretraining on one specific lead: Pretraining exclusively on lead I for 10 epochs improved average AUROC over 12-lead pretraining by 3.4%, 1.4%, and 2.4% for CLEF-S, -M, and -L. This brought CLEF close to the supervised ECGFounder on lead I, with average AUROC differences of only 0.2%, −0.1%, and −0.3% for CLEF-S, -M, and -L. On lead II, improvements were 3.4%, 2.0%, and 0.8%.
  • Ablation table not fully available: The paper references Table 5, "Ablation results on handling missing metadata across CLEF," but its contents are not included in the provided text, so the specific ablation numbers are not reported here.

Methodology in Plain English

Start with a large collection of 12-lead ECGs where no per-recording diagnostic labels exist, but routine patient metadata (such as age, sex, blood pressure, smoking status, diabetes status, and cholesterol) is available. Plug that metadata into an established clinical formula — SCORE2 — which estimates a person's 10-year risk of a cardiovascular event. Every ECG recording now has a risk number attached to it.

Then train a neural network with contrastive learning, the technique that pulls two different augmented views of the same signal together and pushes views of different signals apart. The twist is how "different" is defined. Instead of treating every other sample as equally different, the model is told how different two subjects are according to their risk scores: pairs with very different risk scores get pushed apart strongly, while pairs with similar risk scores are pushed apart only weakly. Two losses work together — a weighted contrastive loss that scales how hard each negative pair is pushed, and a dissimilarity alignment loss that directly nudges the cosine similarity of two embeddings toward the risk-score-derived target.

Because metadata is often missing, the authors add a reliability multiplier that shrinks the weight toward uniform treatment when risk scores are computed from many imputed values. Each ECG sample is randomly assigned one of its 12 leads during pretraining, so one model can later be fine-tuned for any lead. Signals are resampled to 500 Hz over 10 seconds (sequence length 5,000), bandpass filtered between 0.67 and 40 Hz with a Butterworth filter, and z-score normalized. Augmentation randomly selects among muscle noise, movement artifacts, baseline wander, white noise, or no perturbation, each with probability 0.2, using noise derived from free-living ECG recordings. Evaluation uses fine-tuning (all parameters updated) and linear probing (frozen encoder, linear head only), reporting AUROC for classification and mean absolute error (MAE) for regression.

Why This Matters

  • Impact on research: The work shows that clinical domain knowledge, encoded as a validated risk score, can be injected into self-supervised pretraining without any extra labels. It reframes contrastive similarity as a soft, continuous quantity rather than a binary one, and it offers a concrete recipe for handling missing metadata in medical foundation models — a problem that also affects other modalities.
  • Remote and wearable monitoring: Single-lead ECG is the format available in consumer wearables. CLEF maintains performance across leads rather than favoring the lead it was trained on, which matters when the target lead differs by device.
  • Emergency and acute care: The MC-MED results cover emergency department disposition and triage acuity, where fast, single-lead rhythm and arrhythmia assessment can inform triage decisions.
  • Long-term cardiac risk surveillance: The MUSIC task targets long-term outcomes including sudden cardiac death, supporting continuous monitoring beyond a single clinical visit.
  • Blood pressure and cardiac function estimation: The Aurora BP and MC-MED tasks estimate systolic and diastolic blood pressure, and the MIMIC-IV task targets left ventricular ejection fraction, both of which are otherwise invasive or require specialist equipment.
  • Industry relevance: The models are small enough for edge deployment — CLEF-S has only 448K parameters — and the largest variant, CLEF-L at 296M, gives a scaling path for cloud-based analysis. Code and pretrained weights are publicly released, lowering the barrier for device makers and health platforms.

Future Directions

  1. Alternative risk scores and population generalizability. The authors note that SCORE2 was developed on European populations, which may limit generalizability, and that the optimal score may differ by application. Testing the framework with other risk scores and on non-European cohorts is the natural next step.
  2. Closing the gap on specific tasks. Moment outperformed CLEF on MUSIC sudden cardiac death prediction, which the authors attribute to long-range temporal dependencies. Better handling of long-range patterns, or combining contrastive and forecasting-style objectives, is an open direction.
  3. Extending beyond ECG. The framework "supports any validated risk score," which raises the question of whether the same clinically-guided weighting scheme transfers to other physiological signals and multimodal health data.
  4. Improving the smallest model. CLEF-S did not outperform KED or Moment under linear probing and trailed CLEF-M and CLEF-L, so representation quality at small parameter counts remains unresolved for edge deployment.

Target Audience

Researchers and practitioners in machine learning for healthcare, particularly those working on ECG foundation models, wearable biosignal analysis, and self-supervised representation learning. It is also relevant to clinical informatics teams and medical device engineers who want to pretrain models on routinely collected hospital data without per-sample annotations, and to readers interested in how domain knowledge from clinical risk scores can be embedded into contrastive learning objectives.

Authors’ abstract

The electrocardiogram (ECG) is a key diagnostic tool in cardiovascular health. Single-lead ECG recording is integrated into both clinical-grade and consumer wearables. While self-supervised pretraining of foundation models on unlabeled ECGs improves diagnostic performance, existing approaches do not incorporate domain knowledge from clinical metadata. We introduce a novel contrastive learning approach that utilizes an established clinical risk score to adaptively weight negative pairs: clinically-guided contrastive learning. It aligns the similarities of ECG embeddings with clinically meaningful differences between subjects, with an explicit mechanism to handle missing metadata. On 12-lead ECGs from 161K patients in the MIMIC-IV dataset, we pretrain single-lead ECG foundation models at three scales, collectively called CLEF, using only routinely collected metadata without requiring per-sample ECG annotations. We evaluate CLEF on 18 clinical classification and regression tasks across 7 held-out datasets, and benchmark against 5 foundation model baselines and 3 self-supervised algorithms. When pretrained on 12-lead ECG data and tested on lead-I data, CLEF outperforms self-supervised foundation model baselines: the medium-sized CLEF achieves average AUROC improvements of at least 2.6% in classification and average reductions in MAEs of at least 3.2% in regression. Comparing with existing self-supervised learning algorithms, CLEF improves the average AUROC by at least 1.8%. Moreover, when pretrained only on lead-I data for classification tasks, CLEF performs comparably to the state-of-the-art ECGFounder, which was trained in a supervised manner. Overall, CLEF enables more accurate and scalable single-lead ECG analysis, advancing remote health monitoring. Code and pretrained CLEF models are available at: github.com/Nokia-Bell-Labs/ecg-foundation-model.

Read the original paper