Skip to content
AI.info

The Pulse

RADAR Scores 0.913 AUC Across 146 Abdominal CT Findings

A Science study evaluates RADAR, a generalist AI model for abdominal CT diagnosis, across 146 findings and multiple clinical settings.

RADAR Scores 0.913 AUC Across 146 Abdominal CT Findings

AI.info Team ·

RADAR achieves a mean area under the curve of 0.913 across 146 abdominal CT findings, according to a study published September 17, 2026, in Science. The score compares with 0.776 for the best competing vision-language model, giving the system a wide margin across tests covering multiple diseases, organs and clinical settings.

Developed by researchers led by Qi Zhang, RADAR is designed to interpret contrast-enhanced abdominal CT scans as a broad diagnostic system rather than as a tool for one disease or one organ. The researchers describe it as Rapid Abdominal Diagnosis with AI and Radiology. Their results place the model among the strongest reported attempts to build a general-purpose AI reader for one of radiology’s more demanding imaging tasks.

The findings come from a peer-reviewed paper titled “An expert-level generalist AI for abdominal CT diagnosis,” published in Science.

RADAR’s 424,911-exam training base

RADAR was trained on 424,911 examinations containing 1.5 million image-text pairs and more than 15 million anatomy-specific pairs. The training approach divides CT volumes into individual anatomical structures and connects those structures with descriptions in radiology reports.

That design targets a problem specific to abdominal imaging: useful diagnostic signals can be sparse, while the same scan may contain dozens of organs and many possible abnormalities. A model must connect a small visual feature to the correct anatomical location and the relevant clinical description. RADAR uses large-scale contrastive learning to improve those image-report associations.

The study compares RADAR with specialist AI systems and vision-language models across internal and external evaluations. Its headline result is the mean AUC of 0.913 across 146 findings, while the best competing vision-language model reaches 0.776. AUC measures how well a system separates positive and negative cases across a diagnostic task, with higher values indicating better discrimination.

Emergency scans test performance outside training conditions

RADAR also performs strongly on emergency examinations, even though emergency data was not used specifically to train the system. Across more than 27,000 emergency CT cases, the model records an AUC of 0.904.

That result matters because emergency imaging can differ from routine hospital examinations in patient condition, scan timing and the range of suspected disease. The release says RADAR maintained strong performance across organs, diseases and emergency cases, including uncommon findings.

The emergency evaluation does not establish that RADAR can independently manage acute-care diagnosis. It shows that the model’s measured discrimination held up in a large emergency cohort that was not a dedicated part of its training process.

Eight external centers measure generalization

In testing at eight external centers, RADAR achieved an AUC of 0.895. The external evaluation covered patients, clinical settings and imaging protocols beyond the data used to build the model.

The gap between the overall reported score of 0.913 and the external-center score of 0.895 is consistent with the difficulty of moving from development data to new institutions. Differences in scanners, protocols, patient populations and reporting practices can affect medical imaging systems even when the underlying diseases are similar.

RADAR’s results therefore support a claim about performance across the study’s tested settings, not a guarantee of accuracy in every hospital. The paper’s reported evidence covers a broad set of evaluations, while clinical deployment would require additional validation, workflow testing and oversight.

A generalist model still faces the clinical handoff

Radiologists do more than classify findings. They select the relevant clinical context, weigh competing explanations, decide which abnormalities require action and communicate uncertainty in a report. RADAR’s evaluation measures diagnostic discrimination across labeled findings, which is narrower than proving that the system can replace a complete radiology workflow.

The study does show why generalist systems attract attention in abdominal CT. A collection of narrow tools may identify individual conditions, but a broad model could examine the same scan for many findings in one pass. That could be useful when abnormalities fall outside the indication that prompted the scan or when several organs need review together.

The immediate evidence is numerical: 0.913 across 146 findings, 0.904 across more than 27,000 emergency cases and 0.895 across eight external centers. Those results establish the scope of RADAR’s reported testing; they do not by themselves settle how the model should be integrated into patient care.

Source

EurekAlert

Explore

More articles