Skip to content
AI.info

Research

Performance Evaluation of Transfer Learning Based Medical Image Classification Techniques for Disease Detection

Overview Research area: Medical image classification using deep learning — specifically, transfer learning (TL) with pre-trained convolutional neural networks for disease detection in chest X-rays. Te

Performance Evaluation of Transfer Learning Based Medical Image Classification Techniques for Disease Detection
arXiv
2512.04397
Published
2025-12-04
Authors
Zeeshan Ahmad, Shudi Bao, Meng Chen

AI summary

Overview

Research area: Medical image classification using deep learning — specifically, transfer learning (TL) with pre-trained convolutional neural networks for disease detection in chest X-rays.

Technical level: Intermediate. The paper assumes familiarity with convolutional neural networks, transfer learning, and standard classification metrics, but its experimental setup and conclusions are readable without deep mathematical background.

One-sentence scope: The paper benchmarks six pre-trained CNN architectures (AlexNet, VGG16, ResNet18, ResNet34, ResNet50, and InceptionV3) as transfer-learning feature extractors on a custom chest X-ray dataset, comparing them on accuracy metrics, uncertainty, and runtime.

What This Paper Is About

Training a large deep learning model from scratch requires massive amounts of labeled data and compute, which is usually infeasible in medical imaging where labeled datasets are scarce. Transfer learning sidesteps this by reusing a model pre-trained on a large natural-image dataset (here, ImageNet) and retraining only part of it for the new medical task. The paper's goal is to determine, through a systematic comparison, which pre-trained architecture best classifies chest X-ray images of pneumonia, tuberculosis, normal, and "unknown" conditions, and to weigh that accuracy against robustness and computational cost.

Key Contributions

  1. A comparative evaluation of six widely used pre-trained CNN models — AlexNet, VGG16, ResNet18, ResNet34, ResNet50, and InceptionV3 — under a single, consistent transfer-learning pipeline applied to a custom chest X-ray dataset containing normal, pneumonia, tuberculosis, and unknown classes.
  2. An uncertainty analysis performed over multiple random train-validation-test splits, reporting mean loss and standard deviation of loss to assess how consistently each architecture performs, not just how accurately.
  3. A runtime comparison that quantifies the trade-off between classification performance and computational cost, including the finding that InceptionV3 requires 711 seconds while ResNet50 takes nearly three times as long as ResNet18.
  4. A demonstration that, once a well-trained frozen feature extractor is available, a lightweight feedforward network is sufficient for prediction — the classifier head used across all models follows a hidden-node sequence of 512-128-64-64-2.

Main Findings

  • InceptionV3 leads on every metric: On the test set it achieves precision 0.987, recall 0.981, accuracy 0.989, and F1 score 0.985. On the training set it reaches precision 0.994, recall 0.989, accuracy 0.99, and F1 0.991.
  • Deeper ResNets perform better: ResNet50 achieves test precision 0.980, recall 0.976, accuracy 0.970, and F1 0.980. ResNet34 reaches test precision 0.977 and F1 0.970, while ResNet18 attains test precision 0.961 and F1 0.952 — a consistent improvement with depth.
  • VGG16 and AlexNet trail the field: VGG16 reaches test precision 0.935, recall 0.930, accuracy 0.941, and F1 0.939. AlexNet reaches test precision 0.922, recall 0.918, accuracy 0.926, and F1 0.927. Both are described as performing "reasonably well but with lower accuracy."
  • InceptionV3 is also the most consistent: It shows the lowest mean loss (0.0067) and lowest standard deviation of loss (0.0004) in the uncertainty analysis. AlexNet and VGG16 have higher mean losses but relatively low standard deviations, indicating consistent although weaker performance.
  • Accuracy costs time: InceptionV3 requires the longest runtime at 711 seconds. The ResNet family shows a clear correlation between depth and runtime, with ResNet50 taking nearly three times as long as ResNet18.
  • Model size varies substantially: ResNet-18, ResNet-34, ResNet-50, and InceptionV3 have 11.18M, 21.3M, 23.5M, and 25.12M parameters, respectively.
  • Transfer learning helps, but not uniformly: The paper concludes that TL is beneficial in most cases, especially with limited data, but the extent of improvement depends on model architecture, dataset size, and the domain similarity between source and target tasks.
  • A light classifier head suffices: With a well-trained feature extractor, only a lightweight feedforward model is needed to provide efficient prediction.
  • Vision Transformers were excluded: The authors state that ViT-based models have fundamentally different architectures from convolutional models and are therefore outside the scope of this comparative analysis.

Methodology in Plain English

The authors took a publicly available chest X-ray dataset from Kaggle and split it, after shuffling, into 13,028 training images, 761 validation images, and 1,527 test images. The training set contains normal (4,667), pneumonia (3,633), tuberculosis (3,573), and unknown (1,155) images; the unknown class holds images with anomalies that do not clearly represent a known disease state, and was kept as a separate class rather than discarded. Original images of shape (1272, 1040) were resized to (224, 224) and converted to grayscale (one channel), with Random Horizontal Flip, Random Rotation, and Normalization applied for robustness.

For each of the six models, the initial layers of the ImageNet-pretrained feature extractor were frozen and the final set of layers was trained — the assumption being that early layers capture generic image features. Training used the Adam optimizer with a learning rate of 0.001 and a batch size of 32. ResNet models used a kernel size of (3, 3) in the core feature extraction block and (7, 7) in global average pooling layers; Inception building blocks used a kernel size of (3, 3) and strides of (2, 2). Pre-trained weights were stored as ".pt" files in PyTorch. Features from the frozen extractor were aggregated by a global average pooling layer and passed through a feedforward network with hidden nodes in the sequence 512-128-64-64-2. Implementation was in Python with PyTorch, NumPy, Pandas, and Matplotlib, trained on Google Colab with an NVIDIA T4 GPU.

Evaluation used Precision, Recall, Accuracy, and F1 Score derived from the confusion matrix. For uncertainty, the authors repeated training and validation across multiple random splits and recorded the mean and standard deviation of the validation loss. Runtime was measured during validation. The paper reports the classifier output dimension as 2 in the feedforward sequence while describing four image categories; it does not explain this difference.

Why This Matters

The study gives practitioners a concrete, side-by-side comparison of transfer-learning backbones on a medical imaging task, along with the accuracy-versus-compute trade-off — information that is often missing when models are reported in isolation. It reinforces that transfer learning can substitute for expensive training from scratch when labeled medical data is scarce, and it shows that robustness (measured through loss variability) and runtime deserve a place alongside accuracy when choosing a model.

Real-world applications:

  • Pneumonia detection from chest X-rays, where the model can flag white patches in lung fields characteristic of fluid accumulation.
  • Tuberculosis screening in settings where radiologist time is limited and automated pre-reading could prioritize cases.
  • Anomaly flagging through the "unknown" class, helping systems distinguish unusual or poorly represented conditions from normal and known disease states rather than forcing them into an existing label.
  • Resource-constrained deployment, where lighter models such as AlexNet or VGG16 may be preferred over more accurate but slower models like InceptionV3.

Industry relevance: the findings are directly useful for medical imaging software vendors and hospital IT teams deciding which architecture to deploy, since the reported runtimes (for example, 711 seconds for InceptionV3 versus much shorter runs for simpler models) and parameter counts (11.18M to 25.12M) map onto real hardware and latency budgets. The result that a lightweight feedforward head suffices once the feature extractor is trained also suggests a practical deployment pattern where heavy feature extraction is done once and classification is cheap.

Future Directions

  • Optimizing complex models to achieve faster inference with minimal accuracy loss, so that high-performing architectures like InceptionV3 become viable in resource-constrained environments.
  • Extending the comparison to Vision Transformer based models, which the authors explicitly excluded because their architecture differs fundamentally from convolutional models.
  • Investigating further how dataset size and the domain similarity between the ImageNet source task and the medical target task shape the benefit of transfer learning.
  • Refining the lightweight classifier design, since the paper shows a well-trained frozen feature extractor combined with a simple feedforward head can already deliver efficient predictions.

Target Audience

This paper is most useful to machine learning researchers and engineers working on medical image classification, particularly those selecting a pre-trained backbone for a transfer-learning pipeline. It also benefits applied data scientists and clinical informatics teams who need to balance accuracy against runtime and hardware constraints, and students or newcomers to medical imaging who want a clear, metric-driven example of how standard CNN architectures compare on a real chest X-ray classification task.

Authors’ abstract

Medical image classification plays an increasingly vital role in identifying various diseases by classifying medical images, such as X-rays, MRIs and CT scans, into different categories based on their features. In recent years, deep learning techniques have attracted significant attention in medical image classification. However, it is usually infeasible to train an entire large deep learning model from scratch. To address this issue, one of the solutions is the transfer learning (TL) technique, where a pre-trained model is reused for a new task. In this paper, we present a comprehensive analysis of TL techniques for medical image classification using deep convolutional neural networks. We evaluate six pre-trained models (AlexNet, VGG16, ResNet18, ResNet34, ResNet50, and InceptionV3) on a custom chest X-ray dataset for disease detection. The experimental results demonstrate that InceptionV3 consistently outperforms other models across all the standard metrics. The ResNet family shows progressively better performance with increasing depth, whereas VGG16 and AlexNet perform reasonably well but with lower accuracy. In addition, we also conduct uncertainty analysis and runtime comparison to assess the robustness and computational efficiency of these models. Our findings reveal that TL is beneficial in most cases, especially with limited data, but the extent of improvement depends on several factors such as model architecture, dataset size, and domain similarity between source and target tasks. Moreover, we demonstrate that with a well-trained feature extractor, only a lightweight feedforward model is enough to provide efficient prediction. As such, this study contributes to the understanding of TL in medical image classification, and provides insights for selecting appropriate models based on specific requirements.

Read the original paper