Skip to content
AI.info

Research

DermAI: Clinical dermatology acquisition through quality-driven image collection for AI classification in mobile

Overview Research area: Medical computer vision — specifically, smartphone-based skin lesion image acquisition and AI classification for clinical dermatology. Technical level: Intermediate. The paper

DermAI: Clinical dermatology acquisition through quality-driven image collection for AI classification in mobile
arXiv
2511.10367
Published
2025-11-13
Authors
Thales Bezerra, Emanoel Thyago, Kelvin Cunha, Rodrigo Abreu, Fábio Papais, Francisco Mauro, Natália Lopes, Érico Medeiros, Jéssica Guido, Shirley Cruz, Paulo Borba, Tsang Ing Ren

AI summary

Overview

Research area: Medical computer vision — specifically, smartphone-based skin lesion image acquisition and AI classification for clinical dermatology.

Technical level: Intermediate. The paper assumes familiarity with convolutional neural networks, transfer learning, and cross-dataset evaluation, but its central argument is about data quality and acquisition design rather than novel model architecture.

Scope: The paper introduces DermAI, a smartphone application that standardizes dermatological image capture with real-time on-device quality checks, describes the clinical dataset built with it, and benchmarks how well CNN classifiers trained on public data transfer to this new data.

What This Paper Is About

Public dermatology datasets are often limited in skin tone diversity, contain acquisition artifacts such as ink markings and blur, and lack clear imaging protocols, so models trained on them do not hold up in real clinical settings — particularly in public health systems serving varied populations. The authors built DermAI, a lightweight mobile app that guides clinicians and non-specialists through standardized, quality-checked lesion photography during routine consultations, producing data designed from the outset for machine learning training. They then tested whether classifiers trained on existing public datasets generalize to this newly collected clinical data, and whether retraining on it helps.

Key Contributions

  1. A smartphone-based clinical acquisition platform. DermAI runs on standard mobile devices and performs real-time, on-device image quality assessment, prompting users to recapture poor images rather than submitting them, plus local model adaptation. Unlike prior dermoscopy-focused tools, it targets routine clinical capture with ordinary phone cameras.

  2. A new clinical dermatology dataset. Collected over an ongoing two-year period at the Hospital das Clínicas, Universidade Federal de Pernambuco (Brazil), it currently comprises 3,401 images across 200 uniquely annotated lesions (2,273 benign, 608 malignant, 520 pre-malignant), spanning a wide range of skin tones, ethnicities, devices, and environmental conditions. It is to be made available for research on request to the corresponding authors.

  3. A structured annotation and feedback workflow. Each case carries a unique record identifier linking images to metadata: patient age, Fitzpatrick skin phototype, gender, lesion location, lesion mask, device model/OS/camera specifications, and short clinical descriptions producing image-text annotations. Dermatologists mark lesion centers and define circular regions of interest, and all samples are reviewed and validated by the attending dermatologist.

  4. A cross-dataset benchmarking study with ensemble inference. Multiple CNN backbones were trained on the public PAD-UFES-20 dataset and on DermAI data, then evaluated across PAD-UFES-20, DDI, and DermAI, with majority-vote and learned-fusion ensemble strategies compared.

Main Findings

  • Public-data models do not transfer well. Models trained on PAD-UFES-20 and evaluated on DermAI reached a mean accuracy of 0.4975 across architectures, and performed poorly on DDI (mean accuracy 0.1673), with recall especially weak — a concern because recall matters most for avoiding missed malignant cases.

  • Training on DermAI data improves cross-dataset performance. Architectures retrained on DermAI reached a mean accuracy of 0.8035 on DermAI and 0.4150 on DDI, versus 0.1673 on DDI for PAD-trained models.

  • Quality filtering of existing data helps even when it removes samples. Applying DermAI-based quality filtering to PAD (PAD_filter) and retraining improved performance across backbones despite using fewer samples, raising the mean PAD-UFES-20 accuracy from 0.7457 to 0.7817.

  • The best combination is filtered public data plus local data. The mean across architectures trained on PAD_filter + DermAI produced the best overall results: accuracy 0.8505 on PAD-UFES-20, 0.5698 on DDI, and 0.8511 on DermAI.

  • Adding DDI samples to training did not help. The authors attribute this to DDI's scarcity and noise, concluding that data quantity alone is insufficient and that curated acquisition is essential.

  • Ensembles outperform single backbones. For example, on the DermAI test set, Ensemble-fusion trained on DermAI reached accuracy 0.8467, versus 0.8209 for DenseNet (DermAI) and 0.7834 for MobileNet v3 (DermAI).

  • Ensemble strategy affects stability. Majority vote emphasizes consensus among classifiers to reduce classifier-specific variance, while learned fusion trains a small MLP on concatenated model outputs to exploit correlations between confidence patterns; both are aimed at improving predictive stability and sensitivity for malignant-suspect cases.

  • The on-device quality model is lightweight. It uses a MobileNetV3 backbone followed by a linear layer with SiLU activation, dropout of 0.2, and a sigmoid output, producing 4 binary indicators for sharpness, blur, exposure (over/under), and compression artifacts, trained with synthetically distorted smartphone images.

Methodology in Plain English

Clinicians and medical students captured lesion images during routine consultations at a Brazilian public university hospital using smartphone cameras, following acquisition guidelines: no pre-marking, centering the lesion with on-screen guides, and shooting at a standardized distance of approximately 5 cm. Medical students handled image acquisition only; the attending dermatologist reviewed and validated every sample, and the preliminary diagnosis reflects that clinician's judgment. Each image was linked by a unique patient record identifier to a structured form capturing age, Fitzpatrick skin phototype, gender, lesion location, lesion mask, and device attributes, plus a short clinical description.

To control common problems, the app crops a central region and enforces a square aspect ratio, showing a preview of that cropped area during capture; both original and cropped images are stored. A MobileNetV3-based quality model flags sharpness, blur, exposure, and compression artifacts and asks the user to recapture rather than trying to fix the image, keeping badly degraded samples out of the dataset. Lesions flagged as malignant-suspect trigger biopsy follow-up so the dataset can later be updated with histopathological confirmation.

For classification, the authors trained several CNN backbones with average pooling and a fully connected layer using categorical cross-entropy, plus ensemble strategies. Since PAD-UFES-20 has no official splits, they fixed training and test partitions at 80/10/10 and evaluated models across three datasets on accuracy, recall, precision, and F1-score, restricted to the six lesion classes PAD provides (melanoma, nevus, seborrheic keratosis, actinic keratosis, basal cell carcinoma, squamous cell carcinoma) so the three datasets could be compared directly. The focus was on lightweight models suitable for on-device execution.

Why This Matters

The paper argues that because clinical algorithms face substantial domain shift, cross-validation across heterogeneous data sources exposes a weakness that makes such models untrustworthy in real clinical use. It positions DermAI as a model-agnostic, framework-compatible acquisition layer that makes dataset construction itself part of the machine learning pipeline, so datasets can grow over time and support continuous model refinement.

Real-world applications:

  • Primary care triage. General practitioners and non-specialists receive a binary risk assessment (benign versus potentially malignant), supporting early triage where dermatologists are scarce.
  • Specialist decision support. Dermatologists receive a preliminary multi-class suggestion and can confirm, disagree, or indicate uncertainty, entering their own diagnostic hypothesis — generating a feedback loop for monitoring performance in real use.
  • Public health systems serving diverse populations. Standardized capture on ordinary smartphones suits low-resource settings where non-specialist users rely on standard mobile devices, and unsupervised environments are accepted as part of clinical variability rather than excluded.
  • Dataset construction for regulated and research use. The workflow prioritizes biopsy verification for malignant-suspect lesions and supports later label updates with histopathological results.

Industry relevance: the work is directly relevant to developers of mobile medical imaging tools and telehealth platforms, since it demonstrates on-device quality gating and lightweight inference as practical design choices, and to anyone building regulated clinical datasets where acquisition standards and audit trails of annotations matter. It also supports the argument that curated data quality, not raw volume, drives generalization.

Future Directions

  • Dataset growth and longitudinal validation. The authors state the dataset is expected to grow over time and that collection is ongoing; the paper's experiments are described as preliminary, so broader validation across the expanding cohort remains open.
  • Incorporating histopathological confirmation. Labels currently reflect clinical judgment, with malignant-suspect lesions prioritized for biopsy follow-up; systematically updating labels with confirmed pathology is a stated pathway.
  • Improving the hardest transfer case. DDI remained the most challenging dataset due to high noise and variability, and adding DDI samples did not improve training — addressing how to use such noisy, scarce data effectively is unresolved.
  • Closing the accuracy gap between datasets. Even the best configurations score substantially lower on DDI than on PAD-UFES-20 or DermAI, leaving open how far standardized acquisition alone can reduce domain shift.

Target Audience

This paper is most useful to medical computer vision researchers working on dermatology and skin lesion classification, clinical informatics and mobile health engineers who need practical guidance on image acquisition and on-device quality control, and dermatologists or clinical researchers at public health institutions interested in building datasets that support AI training. It is also relevant to regulators and dataset curators evaluating how standardized acquisition protocols and annotation review affect clinical AI reliability. Readers looking for novel network architectures will find the modeling choices conventional; the distinctive contribution is the acquisition and quality-control framework.

Authors’ abstract

AI-based dermatology adoption remains limited by biased datasets, variable image quality, and limited validation. We introduce DermAI, a lightweight, smartphone-based application that enables real-time capture, annotation, and classification of skin lesions during routine consultations. Unlike prior dermoscopy-focused tools, DermAI performs on-device quality checks, and local model adaptation. The DermAI clinical dataset, encompasses a wide range of skin tones, ethinicity and source devices. In preliminary experiments, models trained on public datasets failed to generalize to our samples, while fine-tuning with local data improved performance. These results highlight the importance of standardized, diverse data collection aligned with healthcare needs and oriented to machine learning development.

Read the original paper