Skip to content
AI.info

Research

From ACR O-RADS 2022 to Explainable Deep Learning: Comparative Performance of Expert Radiologists, Convolutional Neural Networks, Vision Transformers, and Fusion Models in Ovarian Masses

Overview Research area: Medical image analysis — computer vision applied to pelvic ultrasound, specifically the comparison of expert human assessment (O-RADS v2022) against deep learning classifiers a

From ACR O-RADS 2022 to Explainable Deep Learning: Comparative Performance of Expert Radiologists, Convolutional Neural Networks, Vision Transformers, and Fusion Models in Ovarian Masses
arXiv
2511.06282
Published
2025-11-09
Authors
Ali Abbasian Ardakani, Afshin Mohammadi, Alisa Mohebbi, Anushya Vijayananthan, Sook Sam Leong, Lim Yi Ting, Mohd Kamil Bin Mohamad Fabell, U Rajendra Acharya, Sepideh Hatamikia

AI summary

Overview

Research area: Medical image analysis — computer vision applied to pelvic ultrasound, specifically the comparison of expert human assessment (O-RADS v2022) against deep learning classifiers and hybrid human-AI systems for characterizing ovarian/adnexal masses.

Technical level: Intermediate. The paper assumes familiarity with ultrasound risk-scoring systems and with common deep learning architectures (CNNs and Vision Transformers), but the abstract presents the comparison at a conceptual level.

Scope (one sentence): A single-center retrospective study comparing 16 deep learning models, radiologist O-RADS v2022 scores, and hybrid human-AI fusion models on 512 adnexal mass images from 227 patients.

What This Paper Is About

The 2022 update to the Ovarian-Adnexal Reporting and Data System (O-RADS) was intended to make ultrasound-based risk stratification more consistent, but human reading of these images still varies between observers, and the system's conservative thresholds can lead to over-calling lesions. Deep learning has separately shown promise for characterizing ovarian lesions. This paper asks how well radiologists actually perform when applying O-RADS v2022, how that compares with leading CNN and Vision Transformer models trained on the same images, and whether combining radiologist scores with model predictions improves diagnosis.

Key Contributions

  1. A direct benchmark of radiologist O-RADS v2022 performance against deep learning models on the same cohort, rather than evaluating either in isolation.
  2. A broad architecture comparison: sixteen deep learning models spanning DenseNets, EfficientNets, ResNets, VGGs, Xception, and Vision Transformers were trained and validated on the same image set.
  3. Hybrid human-AI fusion models: for each deep learning scheme, a model was built that integrates radiologist-assigned O-RADS scores with the model's predicted probabilities.
  4. Evidence that fusion gains are architecture-dependent: combining expert scores with model output significantly helped CNN-based models but did not produce a statistically significant improvement for ViT-based models.

Main Findings

  • Radiologist-only O-RADS v2022 assessment: achieved an AUC of 0.683 and overall accuracy of 68.0%.
  • CNN performance was wide-ranging: AUCs from 0.620 to 0.908 and accuracies from 59.2% to 86.4%, meaning the weaker CNN models performed below the radiologists while the stronger ones substantially exceeded them.
  • Vision Transformer ViT16-384 was the strongest single model: AUC of 0.941 and accuracy of 87.4%, clearly above radiologist-only O-RADS assessment.
  • Hybrid human-AI frameworks significantly improved CNN models, but the improvement for ViT models was not statistically significant (P > 0.05).
  • Authors' overall conclusion: deep learning models markedly outperform radiologist-only O-RADS v2022, and integrating expert scores with AI yields the highest diagnostic accuracy and discrimination. The abstract does not report which specific hybrid configuration produced the single best number, nor does it give per-model hybrid figures.

Methodology in Plain English

The study was a single-center, retrospective cohort analysis. It used 512 adnexal mass images drawn from 227 patients, 110 of whom had at least one malignant cyst. Radiologists applied the O-RADS v2022 criteria to assign risk scores. Separately, sixteen deep learning models from several well-known architecture families — DenseNets, EfficientNets, ResNets, VGGs, Xception, and Vision Transformers — were trained and validated on the images. Finally, for each of those schemes, the researchers built a hybrid model that takes both the radiologist's O-RADS score and the model's predicted probability as inputs, producing a combined prediction. Model performance was then compared against the radiologist-only assessment and against each other. The abstract does not describe the train/test split, the imaging protocol, the reference standard for malignancy, the exact fusion mechanism, or any explainability technique, even though the title references explainable deep learning.

Why This Matters

Impact on research: The paper reframes the evaluation question from "can AI classify ovarian masses?" to "how does AI compare with the current clinical standard, and does combining the two help?" Its finding that fusion benefits CNNs but not ViTs suggests that the value of human-AI collaboration depends on the model architecture, which is a more specific and testable claim than the general assertion that human-AI teams are better.

Real-world applications:

  • Clinical decision support during pelvic ultrasound reading, where a model score accompanies the sonographer's O-RADS assessment.
  • Second-reader or triage support in settings with less subspecialist ultrasound expertise, where O-RADS interpretation is more variable.
  • Reducing false positives and the downstream unnecessary follow-up imaging or surgery that conservative O-RADS thresholds can trigger.
  • Standardizing risk reporting across sites and operators, which the authors explicitly frame as a goal of hybrid paradigms.

Industry relevance: Developers of ultrasound and radiology AI software, PACS and ultrasound vendor integration teams, and groups pursuing regulatory clearance for ovarian lesion characterization tools all have a stake in whether hybrid human-AI workflows are demonstrably better than either component alone — and in the finding that not every architecture benefits equally from human input.

Future Directions

  • External and prospective validation: the abstract describes a single-center, retrospective design, so generalizability to other sites, scanners, and patient populations is untested.
  • Explaining the architecture-dependent fusion effect: understanding why CNNs gained significantly from radiologist scores while ViTs did not would inform how hybrid systems should be designed.
  • Delivering on interpretability: the title promises explainable deep learning, but the abstract does not describe any explainability method or its evaluation; making model reasoning auditable is a natural next step for clinical adoption.
  • Clinical workflow and outcome studies: determining whether improved discrimination translates into fewer unnecessary interventions, better detection of high-risk lesions, and acceptable reader trust in practice.

Target Audience

Radiologists and sonographers who interpret pelvic ultrasound and use O-RADS; medical imaging and computer vision researchers working on classification benchmarks and human-AI collaboration; clinical informatics and regulatory professionals evaluating AI tools for women's imaging; and deep learning practitioners interested in how CNN and Vision Transformer architectures compare on a real clinical task.

Authors’ abstract

Background: The 2022 update of the Ovarian-Adnexal Reporting and Data System (O-RADS) ultrasound classification refines risk stratification for adnexal lesions, yet human interpretation remains subject to variability and conservative thresholds. Concurrently, deep learning (DL) models have demonstrated promise in image-based ovarian lesion characterization. This study evaluates radiologist performance applying O-RADS v2022, compares it to leading convolutional neural network (CNN) and Vision Transformer (ViT) models, and investigates the diagnostic gains achieved by hybrid human-AI frameworks. Methods: In this single-center, retrospective cohort study, a total of 512 adnexal mass images from 227 patients (110 with at least one malignant cyst) were included. Sixteen DL models, including DenseNets, EfficientNets, ResNets, VGGs, Xception, and ViTs, were trained and validated. A hybrid model integrating radiologist O-RADS scores with DL-predicted probabilities was also built for each scheme. Results: Radiologist-only O-RADS assessment achieved an AUC of 0.683 and an overall accuracy of 68.0%. CNN models yielded AUCs of 0.620 to 0.908 and accuracies of 59.2% to 86.4%, while ViT16-384 reached the best performance, with an AUC of 0.941 and an accuracy of 87.4%. Hybrid human-AI frameworks further significantly enhanced the performance of CNN models; however, the improvement for ViT models was not statistically significant (P-value >0.05). Conclusions: DL models markedly outperform radiologist-only O-RADS v2022 assessment, and the integration of expert scores with AI yields the highest diagnostic accuracy and discrimination. Hybrid human-AI paradigms hold substantial potential to standardize pelvic ultrasound interpretation, reduce false positives, and improve detection of high-risk lesions.

Read the original paper