Research
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
Overview Research area: Medical computer vision and multimodal AI — specifically, benchmark datasets for dermatology visual question answering (VQA) and clinical reasoning. Technical level: Intermedia

- arXiv
- 2601.14084
- Published
- 2026-01-20
- Authors
- Abdurrahim Yilmaz, Ozan Erdem, Ece Gokyayla, Ayda Acar, Burc Bugra Dagtas, Dilara Ilhan Erdil, Gulsum Gencoglan, Burak Temelkuran
AI summary
Overview
Research area: Medical computer vision and multimodal AI — specifically, benchmark datasets for dermatology visual question answering (VQA) and clinical reasoning.
Technical level: Intermediate. The paper is a dataset-and-benchmark description rather than a modeling paper; readers benefit from familiarity with VQA benchmarks and dermatologic terminology, but no mathematics or architecture details are required.
Scope: The paper introduces DermaBench, a clinician-annotated dermatology VQA benchmark of 656 clinical images from 570 patients (Fitzpatrick skin types I–VI) carrying roughly 14,474 expert-authored question–answer annotations, released as a metadata-only resource on Harvard Dataverse.
What This Paper Is About
Most public dermatology image datasets support only image-level classification, such as lesion or malignancy recognition, giving one diagnostic label per image and no way to test whether a vision–language model can ground language in fine-grained visual findings. The paper argues that evaluating whether multimodal models can actually describe and reason about skin lesions requires structured, expert-written question–answer supervision rather than scraped or synthetically generated text. To fill that gap, the authors built DermaBench from the Diverse Dermatology Images (DDI) dataset, with every annotation authored and validated by practicing dermatologists.
Key Contributions
- A clinician-authored dermatology VQA benchmark. DermaBench contains 656 images from 570 unique patients spanning Fitzpatrick skin types I–VI, annotated by six expert dermatologists into roughly 14,474 image–question–answer pairs.
- A hierarchical, three-round-reviewed annotation schema. The question set consists of 22 primary questions (Q0–Q21) designed by three expert dermatologists, drawing on dermatologic examination frameworks and standard textbook morphology chapters, and revised across three separate design rounds.
- A consensus and quality-control protocol. Each image was annotated independently by two dermatologists against a pre-agreed consensus checklist, with discrepancies adjudicated in a structured consensus review, followed by a joint dermatologist-and-engineer quality-control pass; 10 images were excluded for unresolved disagreements, insufficient visual quality, or ambiguous clinical content.
- A metadata-only, license-respecting release. Rather than redistributing images, the authors publish only structured metadata on Harvard Dataverse (https://doi.org/10.7910/DVN/Q4LBIW), preserving the upstream DDI licensing restrictions while enabling standardized benchmarking and user-defined train–test splits.
Main Findings
- Dataset scale and demographics: DermaBench comprises 656 images from 570 unique patients, spanning Fitzpatrick skin types I–VI, drawn from DDI, which the authors describe as the only publicly available dermatology image collection designed explicitly for skin tone diversity and fairness research.
- Annotation volume: Annotation with the AnnotatorMed interface produced 14,474 image–question–answer pairs, described in the abstract as approximately 14.474 VQA-style annotations.
- Schema structure: The full schema consists of 22 primary questions (Q0–Q21), organized into three functional blocks — Q0, a "not answered" field; Q1–Q10, fundamental image questions (content type and modality, identifiability and framing, focus, illumination, artifacts, Fitzpatrick skin tone); Q11–Q20, detailed dermatology VQA; and Q21, an integrative summary description.
- Question-type distribution (Table 2): The paper's Table 2 reports 10 single-choice questions (image category, illumination, focus, skin type), 14 multi-select questions (lesion distribution, shape, border, surface, color, lesion types), and 6 open-ended questions (narrative quality notes, reports, explanations, summaries), for a listed total of 30.
- Granular controlled vocabularies: Q13 uses 27 predefined distribution options, Q14 uses 21 morphologic shape categories, Q15 uses 12 border options, Q16 uses 27 surface-feature categories (such as scale, crust, erosion, atrophy, xerosis), Q18 uses 11 chromatic subtypes, and Q19–Q20 encode 38 primary and secondary elementary lesion morphologies.
- Annotation efficiency: AnnotatorMed's dynamic rendering, which shows sub-questions only when a parent option is selected, reduced visual clutter and scrolling time by an estimated 50 percent, with dermatologists completing a full annotation in an average of 3–5 minutes per image.
- Comparison with existing corpora (Table 1): VQA-RAD (radiology, 315 images, 3.5K QA pairs, open-ended), SLAKE (radiology, 642 images, 14K QA pairs, bilingual English/Chinese), PathVQA (pathology, 5K images, 32K QA pairs, multi-type questions), DermaVQA (clinical user-generated content, 3.4K images, 1.5K QA pairs, multilingual, Reddit and IIYI), MM-Skin (various modalities, ~11K images, 27K QA pairs, adds textbook image–text pairs), Lightweight Derm VQA (clinical, 1,038 images, ~7.3K QA pairs, 7 QA per image), and the test set of DermatoLlama (dermoscopic, 210 images, ~21.5k, dermoscopy only). DermaBench is listed as 656 images with ~14.4K QA pairs, six dermatologists, and Fitzpatrick I–VI.
- Cross-dataset compatibility: DermaBench annotation concepts were designed to be broadly compatible with the SkinCon taxonomy, with core morphological attributes aligned at the question level (Q19, Q20, Q16, Q18, Q13); the authors note that expanding the concept set from 48 to 197 concepts produces partial differences in annotation granularity and alignment with SkinCon, and recommend supplying SkinCon annotations alongside the corresponding DermaBench columns for LLM benchmarking.
- Provenance validation inherited from DDI: The DDI images were retrospectively selected from pathology-confirmed cases at Stanford Clinics between 2010–2020; each diagnosis was validated by an expert dermatologist and a dermatopathologist using the corresponding biopsy report, and skin tone labels were assigned using in-person clinical assessments, demographic photographs, and review of clinical images by two expert dermatologists.
- No model evaluation results are reported. The paper introduces the benchmark, the annotation schema, and a standardized VQA prompt for evaluating vision–language models, but it does not report accuracy, baseline scores, or any quantitative model performance results.
Methodology in Plain English
The authors began with an existing, already-validated image collection — the Diverse Dermatology Images (DDI) dataset — chosen because it was purpose-built for skin tone diversity and fairness. Rather than writing new questions ad hoc, three expert dermatologists built a hierarchical question set based on how dermatologists actually examine a lesion: first broad questions about the image itself (modality, framing, focus, lighting, artifacts, Fitzpatrick skin type), then lesion-level questions about count, anatomical location, distribution, shape, border, surface features, size, color, and elementary lesion types, and finally one open-ended summary in which the annotator integrates everything into a short clinical description.
They iterated on this schema across three design rounds, revising clarity, branching logic, and morphological completeness, and produced a consensus guideline defining the meaning of every question and option so different annotators would interpret terms the same way. Six dermatologists then annotated the images through AnnotatorMed, an open-source web annotation interface built within the SCALEMED framework that renders sub-questions only when the relevant parent option is chosen. Each image received independent annotations from two dermatologists; disagreements were resolved in a structured consensus review covering morphology, distribution, diagnosis, and Fitzpatrick assignment, and a final dermatologist-plus-engineer review checked for residual inconsistencies and structural errors. Ten images were dropped. The output is a VQA-style JSON/CSV metadata table with image identifiers, question and answer text, modality tags, Fitzpatrick phototype, diagnostic and morphological labels, question category, annotator identifier, and supporting metadata — with no images redistributed.
Why This Matters
Impact on research. DermaBench targets three gaps the authors identify in prior dermatology multimodal data: the absence of dermatologist-curated ground-truth VQA annotations for clinical dermatology images, reliance on internet-scraped or synthetically generated supervision, and the lack of benchmarks that evaluate visual understanding, language grounding, reasoning, and fairness together. Because the questions are graded from low-level image quality to open-ended morphological reasoning, the benchmark can separate perception failures from reasoning failures, and the Fitzpatrick I–VI coverage makes skin-tone fairness auditing possible rather than aspirational. Releasing metadata only also creates a reusable template for other specialties where image licensing blocks redistribution.
Real-world applications.
- Teledermatology triage: Testing whether a multimodal model can reliably distinguish benign from malignant lesions could support remote triage and reduce delays in dermatologic evaluation, which the authors name as a motivating clinical use case.
- Fairness auditing of clinical AI: The skin-tone distribution and per-image Fitzpatrick labels let developers check whether model errors concentrate in darker skin types, addressing the underrepresentation that the paper links to algorithmic bias.
- Clinical report generation: The open-ended narrative and summary questions (Q21) create a graded way to evaluate model-written clinical descriptions against dermatologist-authored text.
- Dermatology education and reference: The controlled vocabularies for morphology, distribution, border, surface, and color provide structured teaching material and a standardized descriptive lexicon.
Industry relevance. Model builders and vendors developing dermatology-capable vision–language models get a clinically grounded test set they can pair with existing classification benchmarks; because the schema is aligned at the question level with SkinCon, teams already using that taxonomy can add DermaBench evaluation without redesigning their label pipelines. Regulators and health systems interested in evidence of fairness across skin tones have a documented, consensus-driven annotation protocol to point to, and the permissive research release lowers the barrier to third-party reproduction.
Future Directions
- Baseline evaluation is missing. The paper supplies a standardized VQA prompt and recommends evaluation workflows but reports no model results, so the immediate next step is running vision–language models — including dermatology-specific systems — against the 22-question schema.
- Native train–test splits and evaluation tooling. DermaBench deliberately ships no predefined split and includes no code ("There is no code or codebase in this study"), leaving scoring, splitting, and reproducibility infrastructure to downstream researchers.
- Cross-dataset harmonization. The authors flag partial granularity and alignment differences with SkinCon after the concept set grew from 48 to 197 concepts, and the alignment is only at question level — expanding this into a fuller mapping is an open task.
- Broadening modality and demographic coverage. DermaBench is built entirely on DDI; extending expert-verified VQA annotation to other dermatology image collections, additional modalities, and larger patient populations would test whether findings generalize beyond this single source.
Target Audience
Researchers and engineers building or evaluating medical vision–language and multimodal models, particularly those working on dermatology, teledermatology, or clinical report generation. It is also relevant to dermatologists and clinical informaticists interested in structured annotation methodology and skin-tone fairness auditing, and to dataset curators in other medical specialties looking for a metadata-only release model that respects upstream image licensing. Readers seeking model accuracy numbers or training methods will not find them here; the paper is a benchmark construction and validation report.
Authors’ abstract
Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.