Research
A Novel Multi-branch ConvNeXt Architecture for Identifying Subtle Pathological Features in CT Scans
Overview Research area: Medical image analysis / computer vision — deep learning for CT-scan-based disease classification, applied here to COVID-19 diagnosis. Technical level: Intermediate. The paper

- arXiv
- 2510.09107
- Published
- 2025-10-10
- Authors
- Irash Perera, Uthayasanker Thayasivam
AI summary
Overview
Research area: Medical image analysis / computer vision — deep learning for CT-scan-based disease classification, applied here to COVID-19 diagnosis.
Technical level: Intermediate. The paper uses standard deep learning vocabulary (transfer learning, pooling, transfer-learning fine-tuning) but explains each pipeline stage in plain terms.
Scope: The paper proposes a multi-branch ConvNeXt-Small architecture with three parallel pooling branches (global average, global max, and a novel attention-weighted pooling branch), trained on a combined 2,609-slice CT dataset, and reports COVID-19 classification metrics that it states outperform all previously reported models on the same datasets.
What This Paper Is About
COVID-19 CT scans of infected and healthy patients look nearly identical to an untrained eye, and expert review of every scan is slow and requires scarce radiology expertise. The authors set out to build an automated classifier that can reliably detect the subtle pathological features of COVID-19 in chest CT slices, using a modern convolutional backbone rather than the older ResNet/DenseNet-era models that earlier COVID-19 CT studies relied on. Their goal was to combine careful, domain-specific data preparation with a redesigned classification head to reach state-of-the-art performance on a small, imbalanced, two-source dataset.
Key Contributions
-
A novel multi-branch classification head on ConvNeXt-Small. The model feeds backbone feature maps through three parallel paths — Global Average Pooling, Global Max Pooling, and a new Attention-weighted Pooling branch that learns an attention mask multiplied with the base feature maps — then concatenates the three outputs and passes them through a feature-selection dense layer with sigmoid activation before the final classification head.
-
A domain-specific preprocessing pipeline that acts as a manual attention mechanism. Contrast Limited Adaptive Histogram Equalization (CLAHE) is applied to boost local contrast, and each lung is cropped separately, resized to 125x250 pixels, and horizontally concatenated into a single 250x250 image, removing ribs, heart, and background.
-
A two-phase transfer-learning training strategy. Phase 1 trains only the new classification head for 12 epochs at a learning rate of 1×10⁻³ with the ConvNeXt base frozen; Phase 2 unfreezes half of the base model's layers and trains for 8 more epochs at 1×10⁻⁶.
-
A combined two-dataset benchmark with augmentation-based class balancing. The authors merged the COVID-19 CT Lung and Infection Segmentation Dataset (20 labeled COVID-19 CT scans) and the MedSeg Covid Dataset 2 (9 labeled axial volumetric CTs; 373 of 829 slices evaluated as positive and segmented) into 2,609 CT slices, then used augmentation to turn an original 1,724 COVID / 885 non-COVID split into a balanced 2,500-per-class, 5,000-image training set.
Main Findings
-
Final validation metrics: Loss 0.0912, ROC-AUC 0.9937, F1-Score 0.9825, Precision 0.9835, Recall 0.9815, Accuracy 0.9757.
-
Few errors on 701 validation images: the confusion matrix shows only 8 false positives and 9 false negatives out of 701 validation images.
-
AUC and F1 were prioritized over raw accuracy. The authors state they optimized for AUC rather than accuracy because of the original class imbalance, since accuracy can be inflated by predicting the majority class.
-
Improvement over every listed comparator on the same datasets. Compared against CNN 8-layers (Acc 0.7467, Sen 0.8, Spe 0.70, AUC 0.78), InceptionV3 (0.8267, 0.88, 0.78, 0.82), Efficient-Net (0.9067, 0.91, 0.85, 0.93), ResNet+SE (0.8707, 0.9322, 0.8030, 0.9557), ResNet (0.8890, 0.9253, 0.8439, 0.9649), ResNet+CBAM (0.9162, 0.8808, 0.9552, 0.9784), MTL (0.9467, 0.96, 0.92, 0.97), and MA-Net (0.9588, 0.9512, 0.9672, 0.9885). The proposed MB-ConvNeXt reports 0.9757 / 0.9815 / 0.9835 / 0.9937.
-
Training-curve behaviour: validation AUC rose steeply in the first phase while the base model was frozen, then stabilized after fine-tuning began at epoch 12, with validation AUC peaking at 0.9937 and close alignment between training and validation curves.
-
High prediction confidence and class separation: a violin plot of predicted probability distributions shows clear separation between COVID and non-COVID cases, and the ROC curve closely follows the top-left corner.
-
Augmentation removed the need for a specialized loss. Because balancing produced equal class counts, the authors used standard binary cross-entropy with equal class weights instead of Focal Loss, which had been considered during planning.
Methodology in Plain English
The researchers started by loading volumetric CT scans and discarding the uninformative top and bottom of each scan, keeping only slices from 20% to 80% of the volume. Each kept slice was rotated to the correct orientation, resized to 512x512, and normalized to a 0-to-1 range. They then enhanced local contrast with CLAHE so subtle features like ground-glass opacities stand out, and used a masking-and-contour routine to find the two largest lung-shaped regions, crop each lung separately, resize each to 125x250, and stitch them side by side into one 250x250 image.
Because the combined raw data was imbalanced (1,724 COVID vs 885 non-COVID), they applied rotation, horizontal and vertical flipping, shifting, gamma correction, and slight noise to synthesize new minority-class samples, ending with 2,500 images per class and 5,000 training images total. The data was split with a stratified 70/30 train/validation split, giving a 701-image validation set (486 COVID-19, 215 non-COVID-19) after cleaning, with batch size 32.
The classifier sits on top of an ImageNet-pretrained ConvNeXt-Small. Instead of a single pooling step, the feature maps go down three parallel paths — one that averages everything (overall texture and context), one that takes the strongest activations (prominent lesions), and one that learns an attention mask to decide where to look — and the three results are concatenated, filtered through a sigmoid-activation dense layer, and fed to a final classification head ending in a single sigmoid neuron for binary output.
Training proceeded in two phases so that the newly added layers would not destroy the pretrained ImageNet features: first 12 epochs with the backbone frozen at a learning rate of 1×10⁻³, then 8 epochs with half the backbone layers unfrozen at 1×10⁻⁶. Callbacks saved the best weights by validation loss and AUC, reduced the learning rate when validation performance plateaued, and guarded against overfitting.
Why This Matters
Impact on research: The paper argues that older COVID-19 CT classifiers, built on ResNet- and DenseNet-style backbones, were limited partly by the small public datasets available early in the pandemic, and that re-examining these tasks with modern architectures plus disciplined data handling can push results considerably higher. It also provides a template — preprocessing, class balancing, architecture design, staged fine-tuning — that the authors say generalizes to other pathologies in CT scans, not just COVID-19.
Real-world applications:
- Triage support in settings with too few radiologists to review every chest CT.
- Screening where RT-PCR testing is unavailable, since the paper notes CT was shown to be more sensitive than RT-PCR.
- Automated flagging of subtle findings such as ground-glass opacities and consolidation patterns.
- Adaptation of the same pipeline to other CT-based pathology classification tasks.
Industry relevance: The pipeline runs on a widely used convolutional backbone and a modest dataset, and the authors release the source code publicly. Both factors lower the barrier to reproducing or adapting the approach in medical-imaging software, and the paper states the results are competitive with contemporary state-of-the-art benchmarks, which matters for anyone building clinically deployable screening tools.
Future Directions
- Train on larger, more diverse data. The authors state the dataset, while sufficient for a robust proof-of-concept, is still relatively small compared to large-scale clinical settings, and that generalizability could improve with data from multiple hospitals and varied scanner parameters.
- Explore alternative architectures. Future work suggested includes a full Vision Transformer, or a hybrid combining convolutional and transformer layers, to see whether they can achieve even higher performance.
- Validate externally. All comparisons were performed on the same two datasets; whether the model holds up on unseen scanners or patient populations is not reported.
- Clarify the test-set question. The paper labels Table I as evaluation on the "test set" but describes the 701 images as an unseen validation dataset; a fully held-out test evaluation is not reported.
Target Audience
This paper is most useful to graduate students and researchers working on deep learning for medical imaging, especially those interested in transfer learning, pooling strategies, and class-imbalance handling for small CT datasets. Medical-imaging engineers and applied ML practitioners building diagnostic classifiers will find the preprocessing and two-phase training recipe directly transferable. Radiologists and clinical informatics teams evaluating AI screening tools will benefit from the metric discussion, though they should note the limited data scale acknowledged by the authors.
Authors’ abstract
Intelligent analysis of medical imaging plays a crucial role in assisting clinical diagnosis, especially for identifying subtle pathological features. This paper introduces a novel multi-branch ConvNeXt architecture designed specifically for the nuanced challenges of medical image analysis. While applied here to the specific problem of COVID-19 diagnosis, the methodology offers a generalizable framework for classifying a wide range of pathologies from CT scans. The proposed model incorporates a rigorous end-to-end pipeline, from meticulous data preprocessing and augmentation to a disciplined two-phase training strategy that leverages transfer learning effectively. The architecture uniquely integrates features extracted from three parallel branches: Global Average Pooling, Global Max Pooling, and a new Attention-weighted Pooling mechanism. The model was trained and validated on a combined dataset of 2,609 CT slices derived from two distinct datasets. Experimental results demonstrate a superior performance on the validation set, achieving a final ROC-AUC of 0.9937, a validation accuracy of 0.9757, and an F1-score of 0.9825 for COVID-19 cases, outperforming all previously reported models on this dataset. These findings indicate that a modern, multi-branch architecture, coupled with careful data handling, can achieve performance comparable to or exceeding contemporary state-of-the-art models, thereby proving the efficacy of advanced deep learning techniques for robust medical diagnostics.