Skip to content
AI.info

Research

3D CT-Based Coronary Calcium Assessment: A Feature-Driven Machine Learning Framework

Overview Research area: Medical image analysis / computer vision applied to coronary computed tomography angiography (CCTA), comparing hand-crafted radiomics features against deep features from pretra

3D CT-Based Coronary Calcium Assessment: A Feature-Driven Machine Learning Framework
arXiv
2510.25347
Published
2025-10-29
Authors
Ayman Abaid, Gianpiero Guidone, Sara Alsubai, Foziyah Alquahtani, Talha Iqbal, Ruth Sharif, Hesham Elzomor, Emiliano Bianchini, Naeif Almagal, Michael G. Madden, Faisal Sharif, Ihsan Ullah

AI summary

Overview

  • Research area: Medical image analysis / computer vision applied to coronary computed tomography angiography (CCTA), comparing hand-crafted radiomics features against deep features from pretrained medical foundation models for coronary artery calcium (CAC) classification.
  • Technical level: Intermediate. Familiarity with CT imaging, radiomics, and classical machine learning classifiers helps, but the pipeline is described in straightforward terms.
  • Scope: The paper builds a radiomics-based, pseudo-label-driven pipeline for zero versus non-zero calcium score classification from non-contrast CCTA volumes and benchmarks it against CT-FM and RadImageNet feature extractors on a single-centre clinical dataset.

What This Paper Is About

Coronary artery calcium scoring is central to early detection and risk stratification of coronary artery disease, but manual scoring is slow and operator-dependent, and automated deep learning pipelines usually need expert-drawn coronary segmentations that are scarce. This paper asks whether radiomics features, extracted from anatomically relevant regions produced by an automatic segmentation tool rather than by human annotators, can classify patients as having zero or non-zero calcium more effectively than embeddings from pretrained medical imaging foundation models. The goal is an accurate, annotation-free classifier that works on routinely acquired non-contrast CCTA scans.

Key Contributions

  1. A pseudo-labeling / weakly supervised radiomics pipeline. The authors use TotalSegmentator to generate pseudo-segmentations of the left and right coronary arteries directly from full CCTA volumes, providing region-of-interest masks without any human-annotated training data, and then extract PyRadiomics features from those regions.
  2. A head-to-head comparison of three feature extraction strategies. Radiomics (112 features across seven categories, reduced to 36 after correlation thresholding), CT-FM 3D embeddings (512 features), and RadImageNet 2D ResNet50 embeddings are each paired with the same set of classical classifiers under identical evaluation settings.
  3. An analysis of contrast versus non-contrast training data. Two training configurations are compared — a mixed dataset of contrast-enhanced plus non-contrast scans (80% training) and a non-contrast-only dataset (80% training) — with testing always performed on non-contrast scans (20%), reflecting the clinical standard for calcium scoring.
  4. Statistical validation of the feature-type comparison. Paired t-tests quantify whether radiomics models differ significantly from CT-FM models, rather than relying on point estimates alone.

Main Findings

  • Radiomics wins. Radiomics-based models showed the best overall performance, with Random Forest achieving the highest accuracy of 84%, sensitivity of 95%, and specificity of 72% (PPV 0.79, F1-score 0.86, NPV 0.93). The abstract summarizes this as radiomics-based models reaching 84% accuracy at p < 0.05.
  • Other radiomics classifiers were competitive. XGBoost and LightGBM also performed well with accuracies of 81% and 78% respectively in the combined contrast-plus-non-contrast training setting; on non-contrast-only training, Random Forest and LightGBM both reached 0.84 accuracy in Table 2.
  • CT-FM features lagged. Models using CT-FM features showed lower performance, with the paper reporting the MLP model reaching an accuracy of 74% and an F1-score of 82% on the non-contrast dataset. Table 3 lists the non-contrast SVM at 0.74 accuracy with 0.82 F1-score and the non-contrast MLP at 0.69 accuracy with 0.82 F1-score.
  • RadImageNet features performed worst. Classification using RadImageNet features gave the lowest accuracy, with the best model (LightGBM) reaching 63% accuracy on non-contrast data; most models landed between 55% and 60% accuracy with modest sensitivity and specificity.
  • Sensitivity dropped when training only on non-contrast data for CT-FM. The paper notes sensitivity was moderate across CT-FM models and decreased further when trained only on non-contrast data.
  • The difference is statistically significant. Paired t-tests showed significant differences between the radiomics and CT-FM models both on the combined contrast and non-contrast dataset (accuracy p = 0.033; F1-score p = 0.017) and on the non-contrast dataset alone (accuracy p = 0.031; F1-score p = 0.024).
  • Classical tree-based models were the most balanced. Random Forest and XGBoost demonstrated the most balanced and highest performance among the classifiers evaluated.
  • Dataset composition. The internal dataset comprised 188 patients, with CAC scores available for 182 who formed the final analysis cohort: 94 with No Calcification (CAC = 0) and 88 with Calcification (CAC > 0). Clinical reports were available for 185 patients.

Methodology in Plain English

The authors collected an internal cohort of patients who underwent ECG-gated CCTA on a GE Medical Systems Revolution Apex scanner as part of the ACTION registry at the University of Galway. DICOM slices were reconstructed into 3D NIfTI volumes for analysis.

For the radiomics route, they did not ask experts to trace the coronary arteries. Instead they ran TotalSegmentator, a tool that automatically segments major anatomical structures including the left and right coronary arteries, and used its output as a pseudo-segmentation mask. Volumes in which no coronary arteries were detectable were excluded, leaving 181 volumes when training on non-contrast scans only and 485 when training on combined contrast and non-contrast scans. PyRadiomics then extracted 112 features from these regions across seven categories: 14 shape, 18 first-order intensity, 24 GLCM, 14 GLDM, 16 GLRLM, 16 GLSZM, and 5 NGTDM features. Correlation thresholding removed redundant, highly correlated features, leaving a final set of 36.

For the deep feature routes, they used two pretrained models without any fine-tuning. CT-FM is a 3D SegResNet-based model pretrained on 146,000 CT scans from the Imaging Data Commons via label-agnostic contrastive learning, producing 512 features per volume; here volumes were reoriented to the SPL anatomical coordinate system, intensities clipped to [-1024, 2048], linearly scaled to [0, 1], and background regions removed. RadImageNet is a 2D ResNet50 pretrained on 1.35 million annotated medical images covering 165 radiologic labels; it was applied slice by slice on axial images and the per-slice embeddings averaged into one volumetric vector.

Five classical classifiers — SVM, Random Forest, XGBoost, LightGBM, and MLP — were trained on these feature sets. Hyperparameters were tuned by grid search with five-fold cross-validation, and the best configuration of each model was evaluated on an independent non-contrast test set under two training regimes (mixed contrast plus non-contrast, and non-contrast only, both split 80/20). Performance was measured with balanced accuracy, sensitivity, specificity, precision (PPV), F1-score, and NPV. All experiments ran on an NVIDIA RTX 4500 Ada Generator using PyTorch.

Why This Matters

Impact on research. The study pushes back on the assumption that large pretrained foundation model embeddings automatically beat hand-crafted features in small, annotation-poor medical imaging settings. It shows that interpretable, reproducible radiomics descriptors of texture and intensity still provide a more discriminative representation for calcium classification, and that deep feature extractors may lack task-specific sensitivity without fine-tuning or better fusion strategies. It also demonstrates that pseudo-segmentation can substitute for expert annotation, which lowers the barrier for building CAC pipelines.

Real-world applications.

  • Automated early cardiovascular risk stratification from the non-contrast CCTA scans already acquired in routine practice.
  • Triage support for radiologists, flagging patients with non-zero calcium for closer review and preventive treatment planning.
  • Deployment in resource-limited settings, since the approach requires minimal training data and no expert annotation.
  • Integration into clinical workflows alongside existing semi-automated calcium scoring software to reduce the time-consuming, operator-dependent slice-by-slice visual confirmation.

Industry relevance. For medical imaging software vendors and hospital systems, the finding that a lightweight, radiomics-based classifier with pseudo-labels matches or exceeds foundation-model embeddings is directly relevant to product cost, explainability, and regulatory review, since radiomics features are interpretable and the pipeline avoids expensive annotation campaigns.

Future Directions

  • Incorporate clinical report text through multimodal fusion, letting models leverage both visual and textual features.
  • Move beyond binary classification to multi-class calcium scoring — 0 (no), 1–10 (minimal), 11–100 (mild), 101–400 (moderate), and >400 (severe) — for more granular and clinically meaningful risk stratification.
  • Expand the datasets and include segmentation masks to further enhance learned feature extraction.
  • Fine-tune or develop improved fusion strategies for CT-FM and RadImageNet embeddings, since the authors suggest current deep feature extractors may lack task-specific sensitivity.

Target Audience

Researchers and graduate students in medical image analysis, radiomics, and computer-aided diagnosis; clinicians and cardiologists interested in automated calcium scoring; machine learning engineers evaluating foundation models against classical feature engineering in low-annotation regimes; and healthcare technology teams looking for deployable, interpretable risk-stratification tools.

Authors’ abstract

Coronary artery calcium (CAC) scoring plays a crucial role in the early detection and risk stratification of coronary artery disease (CAD). In this study, we focus on non-contrast coronary computed tomography angiography (CCTA) scans, which are commonly used for early calcification detection in clinical settings. To address the challenge of limited annotated data, we propose a radiomics-based pipeline that leverages pseudo-labeling to generate training labels, thereby eliminating the need for expert-defined segmentations. Additionally, we explore the use of pretrained foundation models, specifically CT-FM and RadImageNet, to extract image features, which are then used with traditional classifiers. We compare the performance of these deep learning features with that of radiomics features. Evaluation is conducted on a clinical CCTA dataset comprising 182 patients, where individuals are classified into two groups: zero versus non-zero calcium scores. We further investigate the impact of training on non-contrast datasets versus combined contrast and non-contrast datasets, with testing performed only on non contrast scans. Results show that radiomics-based models significantly outperform CNN-derived embeddings from foundation models (achieving 84% accuracy and p&lt;0.05), despite the unavailability of expert annotations.

Read the original paper