Skip to content
AI.info

Research

VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning

Overview Research area: Medical computer vision and medical visual question answering (Med-VQA), specifically chest radiography, combined with multi-task learning that joins question answering, spatia

VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning
arXiv
2511.00504
Published
2025-11-01
Authors
Dang H. Nguyen, Hieu H. Pham, Hao T. Nguyen, Hieu H. Pham

AI summary

Overview

Research area: Medical computer vision and medical visual question answering (Med-VQA), specifically chest radiography, combined with multi-task learning that joins question answering, spatial lesion localization, and clinical reasoning.

Technical level: Intermediate. The dataset design and the clinical rationale are accessible to a general reader, but the fine-tuning recipe (LoRA, IoU thresholds, a combined VQA plus bounding-box loss) assumes some familiarity with vision-language model training.

Scope: The paper introduces VinDr-CXR-VQA, an openly released chest X-ray dataset of 17,597 question-answer pairs over 4,394 images with radiologist-verified bounding boxes and clinical reasoning text, and benchmarks MedGemma-4B-it on it.

What This Paper Is About

Existing medical VQA datasets either ask and answer clinical questions without saying where a finding is located, or they provide lesion detection annotations without any interactive question-answering. Radiologists need both, since they must be able to check what a model detected and where it localized the finding before trusting an output. The authors build a chest X-ray VQA dataset that supplies answers, bounding boxes, and written clinical reasoning together, then fine-tune a medical vision-language model on it to see whether question answering and spatial grounding can be improved at the same time.

Key Contributions

  1. They introduce VinDr-CXR-VQA, a chest X-ray dataset combining spatial grounding (bounding boxes) with clinical question-answering and reasoning explanations, containing 17,597 question-answer pairs across 4,394 images.
  2. The dataset covers six structured question types (Where, What, Is there, How many, Which, Yes/No), with a balanced question-type distribution (16.4–17.0% each) and a balanced positive/negative sample split (41.7% positive, 58.3% negative) intended to reduce hallucinations on normal cases.
  3. All annotations, including bounding boxes and textual explanations, are curated and verified by board-certified radiologists, with automated verification confirming 100% preservation of expert annotations.
  4. The dataset and evaluation tools are publicly released at huggingface.co/datasets/Dangindev/VinDR-CXR-VQA under a CC BY 4.0 license.

Main Findings

  • VQA performance improves with fine-tuning: Fine-tuning MedGemma-4B-it on VinDr-CXR-VQA raised F1 from 0.558 (pretrained baseline) to 0.624, a +11.8% relative improvement, with gains in Accuracy (0.558 to 0.624), Precision (0.421 to 0.471), and Recall (0.827 to 0.926) on 659 validation images.
  • Spatial grounding becomes possible: The fine-tuned model achieved a mean IoU of 0.615 on true-positive bounding boxes, compared with 0.036 for the pretrained baseline, although detection recall stayed low at 0.090 (precision 0.120, F1 0.103, IoU threshold 0.3).
  • Per-image localization quality rises sharply: "Good" localization (IoU ≥ 0.5) went from 1.1% of validation images (pretrained) to 22.8% (fine-tuned), and "acceptable" localization (IoU ≥ 0.3) from 8.5% to 48.6%. Mean IoU over all predictions rose from 0.036 to 0.133.
  • Multi-task training does not hurt the primary task: Fine-tuning improved the VQA task and simultaneously enabled bounding-box prediction, which the authors present as evidence that VinDr-CXR-VQA works as a multi-modal supervision source.
  • The recall gap is attributed to lesion density: The authors note the low recall (9.0%) reflects a lack of multi-lesion training examples, since validation images average 8.3 lesions each and 54.8% of validation images contain two or more pathologies.
  • Expert review supports annotation quality: Two board-certified radiologists with 8 years of experience reviewed 100 question-answer pairs (0.57% of the data) across reasoning accuracy, answer appropriateness, and spatial reference correctness, with near-perfect inter-rater agreement (κ = 0.89; 95% CI: 0.84, 0.93). The 7 initial disagreements were resolved by consensus and all 100 samples were rated "acceptable."
  • Distinctiveness relative to prior datasets: Among the compared datasets (VQA-RAD, SLAKE, PathVQA, MIMIC-CXR-VQA), the authors state VinDr-CXR-VQA is the only one providing both spatial grounding and explainability.

Methodology in Plain English

The authors started from VinDr-CXR, an existing collection of 18,000 posteroanterior chest X-ray images with expert-validated bounding boxes across 22 thoracic diseases. They kept 14 clinically relevant pathologies with higher prevalence, excluded rare or ambiguous findings, and selected 4,394 images that contained at least one expert-annotated pathology finding.

They then hand-designed prompt templates for six question categories and used Google's Gemini 2.5 Pro vision-language API to generate the natural language parts. For each image-template combination, the API received the X-ray plus its full VinDr-CXR annotation (pathology label and bounding box) and produced a question, an answer that includes a spatial reference in <loc_xmin_ymin_xmax_ymax> format, and a 150-250 word clinical reasoning paragraph. The pathology labels and box coordinates were copied directly from VinDr-CXR without modification, so the model-generated text is checked against verified expert ground truth. Each sample therefore has seven fields: question, answer, reason, type, and difficulty (API-generated), plus gt_finding and gt_location (inherited).

Quality control ran in two layers. Automated scripts validated the structure of all seven JSON fields and cross-checked all 4,394 images against the VinDr-CXR source files, finding zero mismatches. Separately, two radiologists reviewed 100 samples.

For the benchmark, the team fine-tuned MedGemma-4B-it using LoRA (rank r = 16, scaling factor α = 32, dropout 0.1), which added only 0.21% trainable parameters (8.9M) while the base model stayed frozen. Training used AdamW with learning rate 2×10⁻⁵, weight decay 0.01, batch size 4 with gradient accumulation of 4 (effective batch size 16), fp16 mixed precision, and 3 epochs over a 20,880-sample training set, taking roughly 36 hours on a single NVIDIA A6000 GPU (48GB VRAM). The objective combines a VQA cross-entropy loss over answer tokens with an IoU-based bounding-box regression loss, weighted as L_total = L_VQA + λ·L_bbox with λ = 0.5. Boxes are emitted as special tokens embedded in the generated answer, so a single model handles both tasks without architectural changes.

Evaluation used Accuracy, Precision, Recall, and F1 on 659 validation images containing 5,471 ground-truth bounding boxes, plus detection metrics at IoU ≥ 0.3 and per-image localization quality at IoU ≥ 0.5 and IoU ≥ 0.3.

Why This Matters

Impact on research. The paper targets a specific gap: detection datasets such as VinDr-CXR and ChestX-ray14 lack interactive question answering, while text-only VQA benchmarks lack spatial annotations. By releasing a dataset with answer, box, and reasoning together, it gives Med-VQA researchers a testbed where a model's answer can be checked against a location, and it makes the "explainable VQA" claim measurable rather than aspirational. The multi-task result (VQA F1 up, grounding newly enabled) suggests grounding supervision need not come at the cost of answer quality.

Real-world applications:

  • Clinical decision support, where a radiologist reviewing an AI answer can also see the location the model relied on and the reasoning text behind it.
  • Triage and reporting workflows in chest radiography, where structured question types (counting, existence checks, anatomical classification) map onto common reading tasks.
  • Model auditing and safety review, using the balanced 41.7% positive / 58.3% negative composition to test for hallucinations on normal cases.
  • Training and education of radiology trainees, since the 150-250 word reasoning fields articulate differential diagnoses and clinical implications.

Industry relevance. Medical imaging vendors and health AI developers need benchmarks that reward both correctness and verifiability. The paper's fine-tuning setup is deliberately cheap (LoRA adding 8.9M parameters, one GPU), which makes replication feasible for smaller teams. The dataset's permissive CC BY 4.0 release and public hosting on Hugging Face lower the barrier to building and comparing grounded medical VQA systems.

Future Directions

  • Close the lesion-density gap. Training data is not lesion-dense, while validation images average 8.3 bounding boxes and 54.8% contain two or more pathologies; the authors propose incorporating lesion-dense training data to bridge this distribution gap and improve the 9.0% recall.
  • Add structured supervision for multi-instance detection. The paper explicitly calls for structured supervision aimed at detecting and mapping multiple lesions per image, rather than relying on tokens appended to a text answer.
  • Broaden clinical validation. The authors recommend extending evaluation across broader populations and institutions, given that only 100 question-answer pairs (0.57%) received expert review in this study.
  • Enrich lesion diversity and reasoning quality. Future work is framed around greater lesion diversity in the dataset and continued emphasis on explainability, alongside expanding the taxonomy beyond the current six question types.

Target Audience

This paper is most useful to medical computer vision and Med-VQA researchers building or evaluating grounded vision-language systems; to clinical AI teams and regulators who need datasets where AI outputs can be visually and textually verified; to radiologists and clinical collaborators interested in how question taxonomies and reasoning text are constructed and validated; and to graduate students or engineers looking for a concrete, moderately sized, openly licensed benchmark with an end-to-end fine-tuning recipe. Readers without a machine learning background can follow the dataset construction and quality-control sections, but the benchmarking sections require familiarity with detection metrics such as IoU, precision, recall, and F1.

Authors’ abstract

We present VinDr-CXR-VQA, a large-scale chest X-ray dataset for explainable Medical Visual Question Answering (Med-VQA) with spatial grounding. The dataset contains 17,597 question-answer pairs across 4,394 images, each annotated with radiologist-verified bounding boxes and clinical reasoning explanations. Our question taxonomy spans six diagnostic types-Where, What, Is there, How many, Which, and Yes/No-capturing diverse clinical intents. To improve reliability, we construct a balanced distribution of 41.7% positive and 58.3% negative samples, mitigating hallucinations in normal cases. Benchmarking with MedGemma-4B-it demonstrates improved performance (F1 = 0.624, +11.8% over baseline) while enabling lesion localization. VinDr-CXR-VQA aims to advance reproducible and clinically grounded Med-VQA research. The dataset and evaluation tools are publicly available at huggingface.co/datasets/Dangindev/VinDR-CXR-VQA.

Read the original paper