Skip to content
AI.info

Research

ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography

Overview Research area: Medical computer vision / explainable AI — automated ejection fraction (EF) estimation from echocardiography video, using prototype-based (ante-hoc) interpretability rather tha

ProtoEFNet: Dynamic Prototype Learning for Inherently Interpretable Ejection Fraction Estimation in Echocardiography
arXiv
2512.03339
Published
2025-12-03
Authors
Yeganeh Ghamary, Victoria Wu, Hooman Vaseli, Christina Luong, Teresa Tsang, Siavash Bigdeli, Purang Abolmaesumi

AI summary

Overview

  • Research area: Medical computer vision / explainable AI — automated ejection fraction (EF) estimation from echocardiography video, using prototype-based (ante-hoc) interpretability rather than post-hoc explanation.
  • Technical level: Advanced (requires familiarity with prototype neural networks, video feature extractors such as R(2+1)D-18, contrastive/distance-based losses, and cardiac ultrasound).
  • Scope: The paper introduces ProtoEFNet, a video-based prototype-learning model for continuous EF regression that learns dynamic spatio-temporal prototypes and is evaluated on the EchoNet-Dynamic dataset against state-of-the-art black-box and post-hoc explainable models.

What This Paper Is About

Ejection fraction is a key measure of how much blood the left ventricle pumps per beat, and it is normally obtained by manually tracing the ventricle at end-diastole and end-systole — a slow, operator-dependent process with inter-observer variation ranging from 7.6 to 13.9 percent. Deep learning can automate EF estimation, but most models are black boxes, and existing explainability approaches are post-hoc methods (attention weights, gradients) that interpret a prediction only after it is made and do not shape the model's internal reasoning. The goal of this paper is to build an inherently interpretable model that estimates continuous EF from echocardiography video and explains its predictions in clinically meaningful terms.

Key Contributions

  1. Dynamic spatio-temporal prototypes: The paper introduces ProtoEFNet, described as the first video-based prototype model for continuous regression and the first inherently interpretable approach to EF estimation, learning prototypes that capture clinically relevant cardiac motion patterns.
  2. Prototype Angular Separation (PAS) loss: A new loss that enforces discriminative representations across the continuous EF spectrum, increasing separation between prototypes with different EF ranges while preserving the ordinality required for regression.
  3. Accuracy on par with black-box models and state-of-the-art explainability: On EchoNet-Dynamic, ProtoEFNet achieves accuracy comparable to non-interpretable counterparts and outperforms most state-of-the-art models, including the post-hoc explainable model EchoGNN.
  4. Clinically meaningful visualizations: Qualitative analysis shows ProtoEFNet attends to essential cardiac features such as LV wall motion and reduced mitral valve movement in EF < 40% cases, whereas leading black-box methods produce diffuse or clinically irrelevant attention.

Main Findings

  • Quantitative performance: On the EchoNet-Dynamic test set with 64 frames, sampling period 1, and clip-level predictions averaged over the entire video (EV), ProtoEFNet reaches R² = 80, MAE = 4.07, RMSE = 5.47, and F1<40% = 78.
  • Comparison to state of the art: ProtoEFNet outperforms most compared SOTA models, including the post-hoc explainable EchoGNN (R² = 76, MAE = 4.45, F1<40% = 78) and EchoCoTr with the same frame size (1 frame / 36 clips / period 2: R² = 79, MAE = 4.18, RMSE = 5.59; 36 frames / period 4: R² = 81, MAE = 3.98, RMSE = 5.34). Its performance matches Resnet2+1D (R² = 80, MAE = 4.10, RMSE = 5.47, F1<40% = 78).
  • Gap to the top black-box model: CoReEcho reports R² = 82, MAE = 3.90, RMSE = 5.13, F1<40% = 80; the authors state this gap could potentially be narrowed with further training and more extensive hyperparameter tuning.
  • Ablation of the PAS loss: Adding the PAS loss decreases MAE from 4.23 ± 0.11 to 4.20 ± 0.06 and increases F1 from 77.67 ± 2.68 to 79.64 ± 2.10, corresponding to the 2% F1 increase highlighted in the abstract. Standard deviations are computed across 5 repetitions of each experiment on the validation set.
  • Effect of removing other losses: Removing L_PSD improves diversity and sparsity but degrades regression performance (as reflected by F1 and MAE), and some prototypes become outliers with no nearby training samples. Without L_Clst, samples and prototypes are scattered in the embedding space without clear ordinality.
  • Embedding-space problem motivating PAS: Even the most distant prototypes, with EF labels of 15% and 86%, had similar embedding vectors with a cosine similarity of around 0.7 before the PAS loss was introduced.
  • Sparse and faithful explanations: Predictions rely only on prototypes with EF values close to the true label — a prototype with EF of 60% does not influence the prediction for a sample with EF of 13%.
  • Clinically meaningful prototypes: The 60% EF prototype exhibits healthy LV and mitral valve motion, while the 11% EF prototype shows reduced LV motion and a thin LV wall.
  • Better localization than Grad-CAM baselines: Compared with CoReEcho's spatio-temporal Grad-CAM attention, ProtoEFNet focuses on the LV wall and mitral/aortic valves, while CoReEcho and EchoCoTr highlight non-specific or clinically irrelevant regions such as background.
  • Temporal alignment: The model captures the periodic nature of echo videos and aligns spatio-temporal features of the input clip with those of the prototype, assigning a high similarity score when it "looks at" the LV wall during systole in the input and matches it to the same phase in the prototype.

Methodology in Plain English

ProtoEFNet builds on the idea of prototype learning: instead of only learning an abstract decision boundary, the model keeps a set of learnable reference examples ("prototypes"), each tied to an EF value and an importance score. A pre-trained R(2+1)D-18 backbone with a feature module and a Region of Interest (ROI) module extracts spatio-temporal features from a video clip; the ROI module produces occurrence maps that highlight where and when in the clip relevant content appears, and these maps are used as weights for pooling the features.

Each prototype is compared with the pooled features using cosine similarity. A regression layer with m weights combines prototype labels into the final prediction as a weighted average, where the weights come from a softmax over the scaled similarity scores with a small temperature of τ = 0.2, which makes the explanation sparse — dissimilar prototypes get essentially zero contribution.

Training combines several terms: a mean-squared-error regression term, a cluster loss that pulls samples with similar EF labels (within a threshold Δl = 5.0%) toward prototypes, a prototype-sample distance (PSD) loss that ensures each prototype has at least one nearby sample, the new PAS loss that uses angular similarity to push prototypes from different EF regions apart, and an L1 occurrence-map regularizer that uses the LV segmentation mask to penalize activations outside the left ventricle. In the last epoch, each prototype is projected onto — that is, replaced by — the closest training feature whose EF label is within Δl, which makes prototypes visualizable as real cases. Hyperparameters include 40 prototypes, k = 3 in the cluster loss, prototypes initialized randomly with labels uniformly set from 10% to 90%, regression weights initialized to 1, and joint fine-tuning for 30 epochs with batch size 16 on an NVIDIA A100 (40 GB) using PyTorch 2.0.1 and CUDA 12.8.

Why This Matters

  • Impact on research: This work moves interpretability from post-hoc explanation to inherent design for a continuous regression task, showing that prototype learning can be adapted from classification to continuous, ordinal clinical measurements and to video rather than static images.
  • Real-world applications:
    • Automated EF measurement in echocardiography labs to reduce clinician workload and manual tracing time.
    • Screening at scale, where large volumes of echocardiograms can be processed with consistent, transparent outputs.
    • Heart failure flagging via the EF < 40% classification task, described as a strong indicator of heart failure.
    • Clinical decision support and training, where case-based prototype explanations let a clinician see which example-like patterns drove a prediction.
  • Industry relevance: Beyond cardiac ultrasound, the framework generalizes to any continuous, video-based clinical measurement where regulators and clinicians demand transparent reasoning rather than black-box outputs.

Future Directions

  • Addressing prototype learning for uncommon EF ranges, which the authors explicitly name as future work; label imbalance was only handled by oversampling the minority region (EF < 50%) and the authors note that fully addressing imbalance was beyond the scope of the study.
  • Narrowing the performance gap to the leading black-box model (CoReEcho) through further training and more extensive hyperparameter tuning.
  • Extending the dynamic prototype approach to other continuous, video-based clinical measurements beyond ejection fraction.
  • Further study of the trade-off observed in the ablation, where removing the PSD loss improved prototype diversity and sparsity metrics but degraded regression accuracy and left outlier prototypes without nearby samples.

Target Audience

Researchers and practitioners in medical computer vision and explainable AI, cardiac imaging specialists and echocardiographers interested in automated EF assessment, and machine learning engineers working on interpretable models for continuous regression on video data. Readers need some background in deep learning and ideally in cardiac ultrasound to follow the loss formulations and clinical claims.

Authors’ abstract

Ejection fraction (EF) is a crucial metric for assessing cardiac function and diagnosing conditions such as heart failure. Traditionally, EF estimation requires manual tracing and domain expertise, making the process time-consuming and subject to interobserver variability. Most current deep learning methods for EF prediction are black-box models with limited transparency, which reduces clinical trust. Some post-hoc explainability methods have been proposed to interpret the decision-making process after the prediction is made. However, these explanations do not guide the model's internal reasoning and therefore offer limited reliability in clinical applications. To address this, we introduce ProtoEFNet, a novel video-based prototype learning model for continuous EF regression. The model learns dynamic spatiotemporal prototypes that capture clinically meaningful cardiac motion patterns. Additionally, the proposed Prototype Angular Separation (PAS) loss enforces discriminative representations across the continuous EF spectrum. Our experiments on the EchonetDynamic dataset show that ProtoEFNet can achieve accuracy on par with its non-interpretable counterpart while providing clinically relevant insight. The ablation study shows that the proposed loss boosts performance with a 2% increase in F1 score from 77.67$\pm$2.68 to 79.64$\pm$2.10. Our source code is available at: https://github.com/DeepRCL/ProtoEF

Read the original paper