Research
Towards Automated Differential Diagnosis of Skin Diseases Using Deep Learning and Imbalance-Aware Strategies
Overview Research area: Medical image analysis / computer vision — automated classification of skin lesions using deep learning. Technical level: Intermediate. The paper assumes familiarity with convo

- arXiv
- 2601.00286
- Published
- 2026-01-01
- Authors
- Ali Anaissi, Ali Braytee, Weidong Huang, Junaid Akram, Alaa Farhat, Jie Hua
AI summary
Overview
- Research area: Medical image analysis / computer vision — automated classification of skin lesions using deep learning.
- Technical level: Intermediate. The paper assumes familiarity with convolutional networks, Vision Transformers, loss functions, and data augmentation, though it explains its core mechanisms in accessible terms.
- Scope: The paper evaluates a Swin Transformer-based pipeline with imbalance-aware components (BatchFormer, Focal Loss, ReduceLROnPlateau) and several augmentation strategies on the ISIC2019 dermatology benchmark, reporting 87.71% accuracy across eight lesion classes.
What This Paper Is About
Dermatological conditions are increasingly common while access to dermatologists remains limited, and general practitioners often diagnose skin disease with suboptimal accuracy. A major technical obstacle is that public skin lesion datasets are severely imbalanced — some lesion types appear tens of times more often than others — which biases classifiers toward majority classes. The paper's goal is to build a deep learning classifier that performs accurately across all eight ISIC2019 lesion classes by combining a hierarchical Transformer backbone with techniques designed specifically to help under-represented "tail" classes.
Key Contributions
- A hybrid Swin Transformer framework pretrained on ImageNet-1K and modified with a four-stage hierarchical design, combining BatchFormer, Focal Loss, and the ReduceLROnPlateau learning-rate scheduler for imbalanced skin lesion classification.
- A systematic evaluation of segmentation preprocessing, using Meta's Segment Anything Model (SAM) to generate lesion masks across three architectures (ViT, ResNet, Swin Transformer) — and an empirical finding that this preprocessing consistently reduced accuracy.
- A targeted augmentation study applying AutoAugment and Elastic Deformation, with Elastic Deformation applied selectively to disease classes with fewer than 2,000 training samples to directly counter imbalance.
- A broad benchmark comparison on ISIC2019 spanning CNNs (DenseNet121, EfficientNet, ResNet50 + SimAM, U-Net), ensembles, and Transformer models (ViT, Swin Transformer), using a stratified 70% / 15% / 15% train-validation-test split.
Main Findings
- Best configuration: The Swin Transformer combined with BatchFormer, Focal Loss, and ReduceLROnPlateau reached 87.71% accuracy on ISIC2019, the highest reported in the paper.
- Segmentation preprocessing backfired: Applying SAM masks lowered accuracy for every architecture tested — ViT fell from 65.6% to 62.0%, ResNet from 75.6% to 73.0%, and Swin Transformer from 76.6% to 70.8%. The authors attribute this to segmentation artifacts and the loss of subtle diagnostic cues such as color variegation, asymmetric texture, and pigment networks, compounded by SAM not being fine-tuned for dermoscopic images.
- AutoAugment helped CNNs: DenseNet improved from 71.53% to 73.87% with AutoAugment, a gain the paper describes as 2.3%.
- Elastic Deformation helped the Transformer: Swin Transformer rose from 87.34% to 87.71% with Elastic Deformation applied selectively to under-represented classes.
- Class imbalance is severe: The "NV" (melanocytic nevus) class contains more than 12,000 samples, while the "DF" (dermatofibroma) class has fewer than 300.
- Transformer models generally outperformed CNNs. Among CNN-family results: DenseNet121-CE 76.47%, DenseNet121-MWNL 61.37%, DenseNet121 + Fourier Transformer 80.16%, EfficientNet 80.94%, Weighted Ensemble 73.18%, Concatenated Ensemble 68.89%, U-Net 60.52%, ResNet50 + SimAM 69.18%. Among Transformer results: ViT 80.16% and the full Swin Transformer configuration 87.71%.
- A numerical detail in the paper: The Swin Transformer appears with different accuracies across tables — 76.6% in the segmentation comparison (Table 1) and 87.34% in the augmentation comparison (Table 2). The text does not explain the difference in experimental conditions between these two settings.
Methodology in Plain English
The researchers started from a Swin Transformer, a vision model that processes images in local windows and progressively merges them into larger representations, pretrained on ImageNet-1K. They organized it into four hierarchical stages that shrink spatial resolution while increasing feature depth, so the model captures both fine lesion texture and broader shape patterns.
Three additions targeted class imbalance. BatchFormer runs a Transformer encoder across the batch dimension rather than the spatial dimension, so samples in the same mini-batch influence each other's gradients. The authors describe this as a form of virtual data augmentation that lets rare classes borrow gradient signal from common ones. Focal Loss (formula: −αₜ(1−pₜ)^γ log(pₜ)) reduces the loss contribution of confidently classified examples and concentrates learning on hard, misclassified ones, with αₜ weighting classes and γ controlling the focusing strength. ReduceLROnPlateau lowers the learning rate when validation loss stops improving, for finer parameter updates late in training.
On the data side, the ISIC2019 images were split 70/15/15 using stratified sampling to preserve class proportions. The team tried SAM-generated lesion masks to emphasize lesion boundaries, then tried AutoAugment (policy-based transformation combinations such as rotation and color jitter) and Elastic Deformation (Gaussian-based displacement fields that simulate natural variation in lesion shape), with the latter reserved for classes below 2,000 training samples.
The paper does not report the specific values used for the Focal Loss parameters αₜ and γ, nor training epochs or hardware.
Why This Matters
- Research impact: The paper provides a negative result that is useful to the field — general-purpose segmentation (SAM) as a preprocessing step degraded performance across three architectures, suggesting that domain-specific segmentation or lesion-aware attention is needed instead. It also confirms that Transformer backbones plus imbalance-aware training outperform CNN baselines on this benchmark.
- Clinical triage support: A classifier reaching 87.71% across eight lesion classes could help general practitioners flag suspicious lesions when a dermatologist is unavailable.
- Patient self-assessment: The authors explicitly frame the model as a potential self-assessment aid for patients, which matters in regions where dermatology access is scarce.
- Education and second opinions: Such tools can serve as decision support for clinicians with limited specialized dermatology training.
- Industry relevance: The pipeline uses standard, widely available components (Swin Transformer, Focal Loss, ReduceLROnPlateau) and an open benchmark (ISIC2019), making it directly reproducible for teams building dermatology triage software, telemedicine platforms, or clinical decision-support products.
Future Directions
- Domain-specific fine-tuning: The authors note the model was initialized with general-purpose ImageNet weights that may not suit dermoscopic imagery, and recommend fine-tuning on dermatology-specific datasets.
- Stronger imbalance-aware learning: The paper calls for dynamic class-weighted loss functions or adaptive sampling techniques to improve predictions for under-represented lesion types, noting the Swin Transformer has no built-in imbalance mechanism.
- Better segmentation strategies: Given SAM's failure, the authors suggest domain-specific segmentation tools or integrating lesion-aware attention mechanisms rather than general-purpose segmentation as preprocessing.
- Open evaluation gaps: The paper reports only accuracy — per-class performance, sensitivity, and AUC are not reported for the proposed model, leaving it unclear exactly how well the rarest classes (such as the DF class with fewer than 300 samples) are actually classified.
Target Audience
This paper suits machine learning practitioners and graduate students working on medical image classification, especially those dealing with long-tailed or imbalanced datasets. It is also relevant to clinical informatics researchers and dermatology-adjacent product teams evaluating whether Transformer architectures and imbalance-aware loss functions can support diagnostic triage. Readers looking for rigorous per-class clinical validation or comparisons against dermatologist performance will not find those here; the value lies in the architectural combination, the augmentation comparison, and the cautionary result about segmentation preprocessing.
Authors’ abstract
As dermatological conditions become increasingly common and the availability of dermatologists remains limited, there is a growing need for intelligent tools to support both patients and clinicians in the timely and accurate diagnosis of skin diseases. In this project, we developed a deep learning based model for the classification and diagnosis of skin conditions. By leveraging pretraining on publicly available skin disease image datasets, our model effectively extracted visual features and accurately classified various dermatological cases. Throughout the project, we refined the model architecture, optimized data preprocessing workflows, and applied targeted data augmentation techniques to improve overall performance. The final model, based on the Swin Transformer, achieved a prediction accuracy of 87.71 percent across eight skin lesion classes on the ISIC2019 dataset. These results demonstrate the model's potential as a diagnostic support tool for clinicians and a self assessment aid for patients.