Skip to content
AI.info

Research

A benchmark multimodal oro-dental dataset for large vision-language models

Overview Research area: Medical/clinical computer vision and multimodal machine learning, specifically AI for dentistry (oro-dental healthcare), with a focus on benchmark datasets and large vision-lan

A benchmark multimodal oro-dental dataset for large vision-language models
arXiv
2511.04948
Published
2025-11-07
Authors
Haoxin Lv, Ijazul Haq, Jin Du, Jiaxin Ma, Binnian Zhu, Xiaobing Dang, Chaoan Liang, Ruxu Du, Yingjie Zhang, Muhammad Saqib

AI summary

Overview

Research area: Medical/clinical computer vision and multimodal machine learning, specifically AI for dentistry (oro-dental healthcare), with a focus on benchmark datasets and large vision-language models (LVLMs).

Technical level: Intermediate. The paper combines dataset curation of images, radiographs, and clinical text with fine-tuning of existing vision-language models; readers should be comfortable with multimodal model concepts, though the abstract describes the approach in accessible terms.

Scope: The paper presents a large multimodal oro-dental dataset released as a public benchmark, and demonstrates its usefulness by fine-tuning and evaluating state-of-the-art vision-language models on dental anomaly classification and diagnostic report generation.

What This Paper Is About

Progress in AI for oral healthcare depends on having large, multimodal datasets that reflect the complexity of real clinical practice, and such datasets have been scarce. This paper introduces a multi-year oro-dental dataset combining intraoral photographs, radiographs, and detailed clinical text, and presents it as a benchmark for AI dentistry. The authors' goal is to show that the dataset is useful by fine-tuning existing large vision-language models on it and measuring their performance against unfine-tuned versions and a commercial model.

Key Contributions

  1. A large multimodal oro-dental dataset. 8,775 dental checkups from 4,800 patients, collected over eight years (2018–2025), with patients aged 10 to 90.
  2. Multiple linked data modalities per patient record. 50,000 intraoral images, 8,056 radiographs, and detailed textual records including diagnoses, treatment plans, and follow-up notes.
  3. An ethically collected, annotated benchmark resource. Data were gathered under standard ethical guidelines and annotated specifically for benchmarking purposes.
  4. A demonstration of dataset utility through fine-tuning. Qwen-VL 3B and 7B were fine-tuned and evaluated on two clinical tasks — classifying six oro-dental anomalies and generating complete diagnostic reports from multimodal inputs — and compared against base models and GPT-4o.
  5. Public release. The dataset is made publicly available as a resource for future AI dentistry research.

Main Findings

  • Fine-tuning produced substantial gains: The fine-tuned Qwen-VL 3B and 7B models achieved substantial improvements over both their base (unfine-tuned) counterparts and GPT-4o. The abstract does not report specific numbers, margins, or metrics, so the exact size of these gains is not stated in it.
  • Two tasks were successfully benchmarked: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. The abstract does not name the six anomaly categories, nor does it give performance figures for either task.
  • The dataset's validity was supported by the results: The authors interpret the performance improvements as validating the dataset and confirming its effectiveness for advancing AI-driven oro-dental healthcare.
  • A scalable multi-year collection is feasible: The dataset spans eight years of checkups across a wide patient age range, indicating that longitudinal, multi-modal clinical dental data can be assembled at a scale usable for model development. The abstract does not break down per-year or per-age-group statistics.

Methodology in Plain English

The authors assembled clinical dental records over eight years, linking three kinds of information for each case: intraoral photographs, radiographs, and written clinical text such as diagnoses, treatment plans, and follow-up notes. The records were collected following standard ethical guidelines and annotated so they could serve as a benchmark for evaluating AI systems. To test whether the dataset is actually useful, the team took an existing family of large vision-language models (Qwen-VL, in 3-billion and 7-billion parameter versions) and specialized them on this data. They then set two tasks: identifying one of six oro-dental anomalies from a case, and producing a full diagnostic report from the multimodal inputs. Finally, they compared the specialized models against the original unfine-tuned versions of the same models and against GPT-4o to see whether training on the dataset helped. The abstract reports the direction of the outcome (substantial gains) but not the detailed evaluation protocol, metrics, or numerical results.

Why This Matters

  • Impact on research: Multimodal AI for dentistry has been constrained by the lack of large, ethically collected, annotated datasets. This release provides a shared benchmark with images, radiographs, and clinical text together, which allows different research groups to compare methods on the same footing.
  • Real-world applications:
    • Assisting dentists in recognizing oro-dental anomalies from routine intraoral photographs and radiographs.
    • Automatically drafting diagnostic reports, reducing documentation burden in clinical practice.
    • Supporting screening or triage in settings where specialist dental expertise is limited.
    • Serving as training material for dental education and for teaching models to reason across images and clinical notes.
  • Industry relevance: The work is directly relevant to dental imaging hardware and software vendors, clinical AI product developers, and healthcare providers looking for evidence that vision-language models can be adapted to dental workflows. The public availability of the data and the benchmarking of both open models (Qwen-VL) and a commercial model (GPT-4o) make it a practical reference point for evaluating cost and performance tradeoffs.

Future Directions

  • Detailed reporting of results: Because the abstract only states that gains were substantial, the natural follow-up is closer examination of the actual metrics, per-task performance, and how the 3B and 7B models compare with GPT-4o. The abstract does not provide these details.
  • Extending coverage: Expanding the anomaly set beyond the six currently defined classes, and broadening the dataset across more institutions, imaging devices, and geographic populations to test generalization.
  • Longitudinal modeling: The eight-year span and wide patient age range raise the question of whether models can track disease progression and treatment outcomes over time rather than only describing a single visit.
  • Clinical deployment and safety: Moving from benchmark results to validated clinical use requires studies of reliability, bias across patient groups, and how generated reports should be reviewed by clinicians.

Target Audience

Researchers and graduate students working on medical multimodal learning, dataset construction, or vision-language models; dental and oral-health informatics specialists interested in AI applications; and industry practitioners in dental imaging and clinical AI who need a benchmark for evaluating models on real-world oro-dental data. Clinicians with an interest in AI tools will find the dataset description relevant, though the paper's focus is on benchmarking rather than clinical deployment.

Authors’ abstract

The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.

Read the original paper