Research
Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting
Overview Research area: Clinical natural language processing and multimodal medical AI — specifically generative chest X-ray (CXR) radiology report generation, with a focus on lines and tubes (L&T) re

- arXiv
- 2511.21735
- Published
- 2025-11-21
- Authors
- Harshita Sharma, Maxwell C. Reynolds, Valentina Salvatelli, Anne-Marie G. Sykes, Kelly K. Horst, Anton Schwaighofer, Maximilian Ilse, Olesya Melnichenko, Sam Bond-Taylor, Fernando Pérez-García, Vamshi K. Mugu, Alex Chan, Ceylan Colak, Shelby A. Swartz, Motassem B. Nashawaty, Austin J. Gonzalez, Heather A. Ouellette, Selnur B. Erdal, Beth A. Schueler, Maria T. Wetscherek, Noel Codella, Mohit Jain, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro, Stephanie Hyland, Panos Korfiatis, Ashish Khandelwal, Javier Alvarez-Valle
AI summary
Overview
Research area: Clinical natural language processing and multimodal medical AI — specifically generative chest X-ray (CXR) radiology report generation, with a focus on lines and tubes (L&T) reporting. The paper is listed under Natural Language Processing.
Technical level: Intermediate. The concepts (multimodal large language models, report-generation metrics, reader studies) are accessible to a general technical reader, though familiarity with radiology reporting and model evaluation metrics helps.
Scope (one sentence): The paper introduces MAIRA-X, a multimodal AI model trained on a 3.1-million-study Mayo Clinic dataset for longitudinal CXR report generation covering both clinical findings and lines/tubes, and evaluates it quantitatively and in a nine-radiologist user study against original radiologist-written reports.
What This Paper Is About
Radiologists must write detailed narrative reports from chest X-rays, including repetitive descriptions of lines and tubes — catheters, tubes, and drains whose type, tip location, and change over time must be documented on every study. AI report generation could reduce this workload, but prior models have been evaluated mainly on pathological findings, not on L&T elements, and user studies of AI-generated reports have reported large gaps versus radiologist-written reports.
This paper's goal is to close that gap by training a CXR-specialist model at clinical scale, building a dedicated evaluation framework for L&T reporting, and running a retrospective reader study to test whether AI-generated reports are clinically acceptable drafts.
Key Contributions
-
MAIRA-X, a clinically evaluated multimodal model for longitudinal CXR report generation. Built on the MAIRA-2 architecture, MAIRA-X generates both clinical findings and L&T descriptions, and is described as the first CXR-specialist report generation model trained at this scale — using CXR-MAYO-REPORT-GEN, a large-scale, multi-site, de-identified Mayo Clinic dataset of 3.1 million studies (6 million images, 806k patients, acquired 2007–2023).
-
RAD-LT-EVAL, a novel LLM-based L&T-specific evaluation framework. Developed with radiologists, it scores L&T reporting attributes including type, longitudinal change from prior studies, placement (including incorrect placement), and counts. The paper states this is the first large-scale L&T-specific evaluation of CXR report generation models.
-
Large quantitative gains over a state-of-the-art baseline. MAIRA-X surpasses the public MAIRA-2 baseline by 10 percentage points (pp) or more on ROUGE-L, CheXpert/macro-F1-14, RadFact/logical-F1, L&T-type/macro-F1 and L&T-placement/macro-F1 across three holdout datasets. When continually trained on MIMIC-CXR (one epoch on the training split), MAIRA-X outperforms prior work on the official MIMIC-CXR test split.
-
A first-of-its-kind retrospective user evaluation. Nine radiologists of varying experience blindly reviewed 600 studies (three reviews per case) across two cohorts — a Target Set matching the expected clinical distribution and an L&T Set with upsampled L&T and tip positions — covering both pathological and L&T-specific assessments.
Main Findings
-
Comparable critical error rates. The user study found 3.0% of original reports versus 4.6% of AI-generated reports contained at least one critical error. In the aggregate reader analysis, 97.0% [96.1%–97.7% CI] of original reports and 95.3% [94.4%–96.4% CI] of AI-generated reports contained no critical errors, a statistically significant difference (permutation testing p = 0.0057).
-
Similar acceptable-sentence rates. The abstract reports acceptable sentences at 97.8% for original versus 97.4% for AI-generated reports; the Results section reports error-free sentences at 97.7% [97.4%–98.1% CI] for original reports and 97.4% [97.0%–97.7% CI] for AI-generated reports.
-
The gap in "no changes needed" reports is about 5 pp. 84.5% [82.8%–86.1% CI] of original reports were acceptable as-is versus 79.4% [77.4%–81.2% CI] of AI-generated reports (permutation testing p < 0.0001). The Discussion reports the gaps as 0.3 pp for error-free sentences, 1.7 pp for reports with critical errors, and 5.1 pp for reports requiring no changes.
-
Omissions are the main asymmetry. 8.0% [6.7%–9.3% CI] of original reports had at least one omission versus 12.7% [11.0%–14.2% CI] of AI-generated reports; multiple omissions were rare in both (0.8% [0.4%–1.2% CI] original, 1.3% [0.8%–1.9% CI] AI-generated).
-
Sentence error proportions are close. Sentence errors appeared in 11.0% [9.5%–12.4% CI] of original and 12.3% [10.8%–14.0% CI] of AI-generated reports; multiple sentence errors in 1.8% [1.2%–2.4% CI] and 1.7% [1.1%–2.3% CI] respectively.
-
Strong results on the public MIMIC-CXR test split. MAIRA-X achieved ROUGE-L 41.3 [41.0, 41.6], CheXpert/macro-F1-14 47.2 [46.5, 47.9], CheXpert/micro-F1-14 64.1 [63.6, 64.5], CheXpert/macro-F1-5 53.2 [52.3, 54.0], CheXpert/micro-F1-5 61.8 [61.1, 62.4], RadFact/logical-precision 61.0 [60.6, 61.4] and logical-recall 55.1 [54.7, 55.5]. For comparison, MAIRA-2 scored ROUGE-L 38.4 [37.8, 39.1] and CheXpert/macro-F1-14 42.7 [40.9, 44.4]; Libra 36.2 and 40.2; LLaVA-Rad 30.6 and 39.5; Med-PaLM M 27.29 and 39.83; Med-Gemma 13.0 and 35.8. MAIRA-2's RadFact logical-precision was 52.5 [51.6, 53.5] and logical-recall 48.6 [47.7, 49.6]; the other models' RadFact values are not reported in the table.
-
L&T gains grow with device count. L&T-counts accuracy was higher for MAIRA-X, and the gap over MAIRA-2 increased as the number of L&T in a study rose from zero to three-or-more, with the highest gains (25 pp or more) at three or more lines/tubes — the authors note this suggests particular value in high-volume settings such as the ICU.
-
Incorrect-placement scores are low in absolute terms. Only 8.4% of all L&T in CXR-MAYO-REPORT-GEN are misplaced, and MAIRA-X nonetheless improved over the baseline on this metric; the authors attribute the low absolute F1 to that low prevalence.
-
Improvement over a prior reader study. Compared with Flamingo-CXR as reported in Tanno et al. (2025), MAIRA-X showed 4.6% (±1.0%) reports with at least one critical error versus 18%; 20.6% (±1.9%) reports with any error versus 30%; 0.06 (±.01) critical errors per report versus 0.28; and 0.29 (±.03) errors of any kind per report versus 0.49. The authors describe this as a high-level comparison because the two evaluation datasets differ.
-
Error types skew toward pathology. Using GPT-based classification, roughly 63.3% of flagged errors were pathological and 34.2% were L&T-related overall. In original reports, 58.5% were pathological and 38.3% L&T; in AI-generated reports, 67.2% were pathological and 31.1% L&T (Chi-squared test p = 0.02). Top error categories were Atelectasis (15%), Pleural effusion (12%), Cardiomegaly (10%), ETT tube positioning (9%), PICC line placement (8%), Pulmonary vascular congestion (7%), Calcified aorta (6%), Pleural thickening (5%), Surgical clips (4%), CVC positioning (4%), and all the rest (20%).
-
High inter-rater variability among radiologists. Average Kendall's concordance across reviewers was W = 0.44. Over 80% of flagged sentence errors or omissions were identified by only one of the three reviewers. Of all agreement sentences, 553 of 671 (82.41%) were flagged by a single reviewer and 118 of 671 (17.59%) by multiple reviewers; for original reports, 258 of 302 (85.43%) versus 44 of 302 (14.57%); for AI-generated reports, 295 of 369 (79.95%) versus 74 of 369 (20.05%).
-
Consensus analysis narrows the critical-error gap. When the three most senior radiologists reclassified all cases flagged as critical by at least one reviewer, reports free from critical errors rose from 95.3% to 96.9% for AI-generated reports and from 97.0% to 98.8% for original reports; pathological-finding errors remained the most frequent critical error type in both.
-
Performance is harder on the L&T-heavy cohort. Both radiologists and MAIRA-X performed more strongly on the Target Set than on the L&T Set, indicating images with more lines and tubes are more difficult to analyze.
Methodology in Plain English
The team took an existing CXR-specialist multimodal model, MAIRA-2, as the base architecture and adapted it for longitudinal reporting that covers both findings and L&T. Training used CXR-MAYO-REPORT-GEN, a de-identified Mayo Clinic dataset of about 3.1 million studies (6 million images from 806k subjects, 2007–2023). They fine-tuned the vision encoder and the language model, adjusted hyperparameters and LLM prompts, and designed L&T performance measures for optimization. Model parameters were selected using a 40,000-study validation set.
Quantitative evaluation used three holdout subsets with different L&T distributions: a test set of 40,000 studies, a 300-study Target Set whose L&T distribution mimics the expected clinical setting, and a 300-study L&T Set with upsampled L&T. They also trained the checkpoint for one additional epoch on the MIMIC-CXR training split and evaluated on the official MIMIC-CXR test split, so that comparisons with models trained primarily on MIMIC-CXR were in-domain.
Metrics came in three groups: lexical (ROUGE-L), clinical efficacy (CheXpert macro/micro-F1 and RadFact logical precision/recall), and the new RAD-LT-EVAL L&T metrics (type, change, placement, incorrect placement, and counts). Confidence intervals for these comparisons came from 500 bootstrapped samples.
For the human evaluation, nine radiologists of varying experience each reviewed reports blindly, with three reviews per case across 600 studies drawn from the Target Set and L&T Set. Reviewers could mark sentences as acceptable, flag clinically insignificant or critical errors, propose corrected sentences, and note omissions. Inter-rater agreement was quantified with Kendall's concordance, and a follow-up consensus pass used the three most senior radiologists to reclassify disputed critical errors by majority vote. Confidence intervals in the reader study came from 1,000 bootstrap resamples.
Why This Matters
Impact on research. The paper shifts the evaluation target for CXR report generation from pathology-only scoring toward L&T-specific assessment, supplying a metric framework (RAD-LT-EVAL) and a blinded multi-reader protocol that other groups can reuse. It also demonstrates that training scale and institutional data matter: MAIRA-2, trained on public multi-institution datasets, did not generalize well to the Mayo Clinic dataset, and MAIRA-X improved by 10 pp or more across three holdout sets. The inter-rater results (W = 0.44; over 80% of flagged errors seen by only one of three reviewers) are a caution for anyone interpreting single-reviewer AI error rates.
Real-world applications:
- Draft report generation for high-volume CXR reading, reducing repetitive L&T documentation for radiologists and potentially improving turnaround times and patient safety.
- ICU and emergency department support, where frequent and precise L&T reporting is needed and where MAIRA-X showed its largest count-based gains (25 pp or more at three or more lines/tubes).
- Quality assurance and second-read workflows, where AI drafts could flag missing devices, tip positions, or interval changes against priors.
- Benchmarking and regulatory-style evaluation of generative radiology models, using the L&T metrics and reader-study design.
Industry relevance. The work is a joint Microsoft–Mayo Clinic effort, and the framing centers on clinical deployment readiness: comparable critical error rates (3.0% original vs 4.6% AI) and comparable acceptable sentence rates, contrasted with an 18% critical error rate reported for Flamingo-CXR in Tanno et al. (2025). That framing targets the practical bar vendors and health systems must clear before drafting tools reach production.
Future Directions
- Reducing omissions. AI-generated reports had more reports with at least one omission (12.7% vs 8.0% for original), which is the largest gap in the reader study and an obvious target for further training or prompting work.
- Improving incorrect-placement reporting. The paper notes MAIRA-X struggled more on L&T incorrect-placement scores because of the very low prevalence of misplaced L&T in training data (8.4% of all L&T). Data curation or sampling strategies to address this class imbalance are an open question.
- Extending the consensus and inter-rater analysis. The consensus reclassification with the three most senior radiologists raised critical-error-free rates to 96.9% (AI) and 98.8% (original), and the authors note that accounting for reviewer variability could lead to better measured performance for MAIRA-X; how best to incorporate majority or consensus scoring into future reader studies remains open.
- Generalization beyond the evaluated settings. MAIRA-2's weak transfer to the Mayo dataset raises the question of how MAIRA-X in turn performs at institutions with different scanners, populations, and reporting styles; the paper does not report a prospective evaluation or deployment results, and the supplied content does not include a formal future-work or conclusion section.
Target Audience
This paper is most useful for medical AI and clinical NLP researchers working on report generation and multimodal models; radiologists and radiology informatics staff evaluating AI drafting tools; and teams in healthcare AI product, regulatory, and clinical-deployment roles who need concrete numbers on where AI-generated CXR reports now agree with—and diverge from—radiologist-written reports. Readers interested in human evaluation methodology will also find the inter-rater variability analysis directly relevant.
Authors’ abstract
AI-assisted report generation offers the opportunity to reduce radiologists' workload stemming from expanded screening guidelines, complex cases and workforce shortages, while maintaining diagnostic accuracy. In addition to describing pathological findings in chest X-ray reports, interpreting lines and tubes (L&T) is demanding and repetitive for radiologists, especially with high patient volumes. We introduce MAIRA-X, a clinically evaluated multimodal AI model for longitudinal chest X-ray (CXR) report generation, that encompasses both clinical findings and L&T reporting. Developed using a large-scale, multi-site, longitudinal dataset of 3.1 million studies (comprising 6 million images from 806k patients) from Mayo Clinic, MAIRA-X was evaluated on three holdout datasets and the public MIMIC-CXR dataset, where it significantly improved AI-generated reports over the state of the art on lexical quality, clinical correctness, and L&T-related elements. A novel L&T-specific metrics framework was developed to assess accuracy in reporting attributes such as type, longitudinal change and placement. A first-of-its-kind retrospective user evaluation study was conducted with nine radiologists of varying experience, who blindly reviewed 600 studies from distinct subjects. The user study found comparable rates of critical errors (3.0% for original vs. 4.6% for AI-generated reports) and a similar rate of acceptable sentences (97.8% for original vs. 97.4% for AI-generated reports), marking a significant improvement over prior user studies with larger gaps and higher error rates. Our results suggest that MAIRA-X can effectively assist radiologists, particularly in high-volume clinical settings.