Research
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
Overview Research area: Medical computer vision and vision-language models (VLMs), specifically radiologic visual question answering (VQA) on CT and MRI. Technical level: Intermediate. The paper is re

- arXiv
- 2512.17396
- Published
- 2025-12-19
- Authors
- Léo Butsanets, Charles Corbière, Julien Khlaut, Pierre Manceron, Corentin Dancette
AI summary
Overview
Research area: Medical computer vision and vision-language models (VLMs), specifically radiologic visual question answering (VQA) on CT and MRI.
Technical level: Intermediate. The paper is readable without deep technical background, but familiarity with VLMs, fine-tuning stages (alignment, instruction tuning), and benchmark evaluation metrics helps.
Scope: The paper introduces and evaluates RadImageNet-VQA, a large-scale CT/MRI dataset and benchmark for radiologic VQA, showing that current vision-language models handle anatomy and basic abnormality recognition but still fail at fine-grained pathology identification.
What This Paper Is About
Existing medical VQA datasets are small, dominated by X-ray images or textbook-style biomedical figures, and often contain linguistic shortcuts that let a model guess the right answer without ever looking at the image. The authors build RadImageNet-VQA from expert-annotated CT and MRI exams to provide a much larger, multimodal, shortcut-resistant resource for both training and benchmarking radiologic question answering.
Key Contributions
-
A large-scale radiologic dataset and benchmark. RadImageNet-VQA includes 750K images paired with 7.5M generated samples (750K image-captions for visual-text alignment plus 6.75M QA pairs), spanning 8 anatomical regions, 97 pathology categories, and three task families (abnormality detection, anatomy recognition, pathology identification) in open-ended, closed-ended, and multiple-choice formats. A stratified benchmark of 1,000 images with 9,000 QA pairs is provided separately.
-
A shortcut-minimizing generation pipeline. Questions are produced with task- and format-specific templates (2 to 7 linguistic variations per task and question type), with distractors drawn from other anatomical regions or clinically plausible diseases from the same region, and "no pathology seen" included in every pathology multiple-choice question.
-
Extensive evaluation of state-of-the-art VLMs. General-purpose and medical-oriented models are benchmarked zero-shot, showing strong anatomy recognition but severe weakness in pathology identification, particularly open-ended.
-
Text-only ablation and fine-tuning studies. Removing images causes open-ended accuracy to collapse to near-random levels, confirming the dataset suppresses linguistic shortcuts, while fine-tuning produces consistent gains across model families—and medically pretrained vision encoders (MedSigLIP) yield no advantage over standard SigLIP.
Main Findings
-
Anatomy recognition is close to solved. InternVL3.5-8B reaches 93.3% on anatomy multiple-choice zero-shot, and after fine-tuning anatomy multiple-choice accuracy is 98.5–99.4% across models.
-
Fine-grained pathology identification is the primary bottleneck. Most zero-shot models score below 20% on open-ended pathology questions; the best performer on that task, MedGemma-4b, reaches only 30.6%.
-
General-purpose models outperform medical-oriented ones overall. InternVL3.5-14B achieves the highest average zero-shot accuracy at 63.6%. Medical models lead only in select areas, such as MedGemma's open-ended pathology performance.
-
Proprietary models struggle on abnormality detection. GPT-5 scores 27.5% and Gemini 2.5 Pro 17.8% on abnormality detection, below random guessing; the authors speculate that aggressive safety alignment pushes models toward conservative "no abnormality" answers.
-
Text-only performance collapses. On VQA-RAD and SLAKE, models reach 11.5–33% accuracy without images, while RadImageNet-VQA open-ended accuracy falls to 2–10%. On multiple-choice, MMMU-Med-val stays above random, whereas RadImageNet-VQA collapses to the expected 25% baseline.
-
Fine-tuning produces substantial and consistent gains. All models improve, with average increases of +19.5% to +22.5%. Abnormality detection shows the largest relative gains (+16.4–39.8%, reaching 85.9–87.7%), while pathology remains the bottleneck.
-
Medical pretraining of the vision encoder does not help. LLaVA-OneVision variants using standard SigLIP outperform those using MedSigLIP across VQA-RAD, SLAKE, and RadImageNet-VQA. Lingshu-7B and Qwen2.5-VL-7B converge to nearly identical performance, suggesting radiologic supervision rather than prior medical pretraining drives downstream capability.
-
Sampling strategy matters modestly. Alternating sampling (each batch drawn from a single dataset, alternating sources) yields slightly stronger and more stable results than mixed sampling, particularly early in training.
-
"No pathology seen" option exposes a bias. Varying the frequency of this option at 0%, 30%, and 100% causes medical models to drop in accuracy (MedGemma 60.9 → 57.6 → 50.2; Lingshu-7B 59.2 → 54.9 → 49.1), while general-purpose models remain comparatively stable (InternVL3.5-8B 45.0 → 44.7 → 46.5; Qwen2.5-VL-7B 40.5 → 40.3 → 42.6).
-
The LLM judge is validated against humans. On a stratified subset of 380 examples, the LLM judge matches human evaluation in 90.6% of cases with a Cohen's kappa of 0.72. Reviewer 1 agreed with the judge on 94.20% of anatomy and 81.05% of pathology cases (kappa 0.7273); Reviewer 2 on 93.98% and 77.27% (kappa 0.7135).
-
The benchmark spans both normal and abnormal studies. The 1,000-image benchmark contains 28.1% normal and 71.9% abnormal studies, with anatomy, pathology, and question-type distributions preserved from the test split.
Methodology in Plain English
The authors start from RadImageNet, a large expert-annotated medical imaging resource where each image carries a modality (CT, MRI, or ultrasound), a body part, and a pathology label. They use only the CT and MRI portions. Because raw labels are not natural language, they convert the metadata into structured radiologic captions using varied radiology-aware templates—for example, verbalizing modality, anatomy, and pathology into a sentence-like description.
They then apply a second set of templates to turn each annotated image into question-answer pairs across three tasks. Anatomy recognition asks which region is imaged. Abnormality detection asks whether anything abnormal is present. Pathology identification asks the model to name or confirm a specific disease within a given anatomical context. Each task is instantiated in open-ended, closed-ended, and multiple-choice forms, and closed-ended questions come in matched positive and negative versions so models cannot simply answer "yes" by default.
Distractors for multiple-choice items are chosen deliberately: anatomy distractors come from other regions in the dataset, and pathology distractors come from clinically plausible diseases in the same region, so a model cannot solve the question by matching anatomy alone. The option "no pathology seen" is included in every pathology multiple-choice question.
The full training split yields roughly 750K images and 7.5M samples. For evaluation, they sample 1,000 images from the test split while preserving the distribution of regions, pathologies, and abnormal cases, producing 9,000 QA pairs. Closed-ended and multiple-choice answers are scored by exact match or rule-based parsing of the option letter; open-ended answers are graded by an LLM-as-a-judge framework (Mistral-Large 2.1 within MedEvalKit), which the authors validate against two blinded human reviewers.
Experiments proceed in three parts. First, a zero-shot benchmark of general-purpose and medical-oriented VLMs. Second, a text-only ablation where images are withheld, to measure how much of the performance comes from linguistic priors. Third, fine-tuning runs on a combined corpus—RadImageNet-VQA train split, CT/MRI datasets converted into 2D VQA pairs (KiTS22, AbdomenAtlas), and existing radiologic VQA datasets (VQA-RAD, SLAKE, and the radiology subset of LLaVA-Med)—using a two-phase recipe: an alignment stage that updates only the vision encoder and projection layers with the language model frozen, followed by instruction tuning of the full model. Ablations compare SigLIP against MedSigLIP vision encoders and mixed against alternating data sampling.
Why This Matters
Impact on research. RadImageNet-VQA shifts medical VQA evaluation toward CT and MRI, which are underrepresented in existing benchmarks, and provides a training corpus large enough for genuinely data-hungry multimodal work. Its text-only ablation offers a reusable methodology for auditing whether a benchmark can be gamed without images—an issue the authors show affects VQA-RAD, SLAKE, and MMMU-Med-val.
Real-world applications.
- Radiologist decision support and triage tools that must answer specific questions about a CT or MRI study rather than produce a free-text report.
- Structured model evaluation for regulatory or procurement decisions, where a benchmark that resists text-only shortcuts is more informative than similarity-based report metrics.
- Medical education and training systems that quiz trainees on anatomy and pathology in imaging studies.
- Diagnostic quality assurance, where automated systems flag anatomy or abnormality discrepancies across large study volumes.
Industry relevance. The finding that standard SigLIP beats medically pretrained MedSigLIP for CT/MRI VQA directly informs how medical AI teams allocate pretraining budgets. The result that fine-tuning on radiologic supervision dominates prior medical pretraining suggests that task-aligned data, not encoder pedigree, drives downstream capability. The shortcut analysis also gives a concrete sanity check that vendors and regulators can apply to any proposed medical VQA benchmark.
Future Directions
-
Multi-finding annotations. The current dataset inherits a single-label pathology per image from RadImageNet, which does not reflect clinical cases with multiple co-existing findings; adding multi-finding samples with richer clinical context is a natural extension.
-
Volumetric and cross-modality work. The dataset is 2D, so models lack volumetric context. The authors note that including both CT and MRI enables future cross-modality studies, and that volumetric approaches may benefit from stronger 2D CT/MRI priors.
-
Closing the pathology gap. Fine-grained identification of subtle, localized, or rare findings—such as patella pathology, coalition bone fusion, quadriceps pathology, Lisfranc ligament injury, and ACL tears—remains unsolved even after fine-tuning, pointing to better visual grounding as the key research target.
-
More natural clinical language. Questions are programmatically generated, which improves control and scalability but may under-represent how clinicians actually phrase questions; the automatic judge for open-ended answers is validated but may still carry residual uncertainty on borderline cases.
Target Audience
Researchers and engineers building medical vision-language models, particularly those working on radiology VQA, multimodal instruction tuning, or CT/MRI image understanding. It is also relevant to clinicians and radiologists interested in how current AI systems perform on imaging questions, to benchmark designers concerned with linguistic shortcuts and distractor quality, and to teams deciding whether to invest in domain-specific vision encoders.
Authors’ abstract
In this work, we introduce RadImageNet-VQA, a large-scale dataset designed to advance radiologic visual question answering (VQA) on CT and MRI exams. Existing medical VQA datasets are limited in scale, dominated by X-ray imaging or biomedical illustrations, and often prone to text-based shortcuts. RadImageNet-VQA is built from expert-curated annotations and provides 750K images paired with 7.5M question-answer samples. It covers three key tasks - abnormality detection, anatomy recognition, and pathology identification - spanning eight anatomical regions and 97 pathology categories, and supports open-ended, closed-ended, and multiple-choice questions. Extensive experiments show that state-of-the-art vision-language models still struggle with fine-grained pathology identification, particularly in open-ended settings and even after fine-tuning. Text-only analysis further reveals that model performance collapses to near-random without image inputs, confirming that RadImageNet-VQA is free from linguistic shortcuts. The full dataset and benchmark are publicly available at https://huggingface.co/datasets/raidium/RadImageNet-VQA.