Computer vision
Medical Imaging Systems
Study medical imaging modalities, labels, leakage, external validation, reader interaction, regulation, monitoring, and safe deployment.
By the end you can
- Explain how modality, protocol, population, and clinical workflow shape medical imaging tasks
- Identify leakage, label, prevalence, and site-shift risks
- Distinguish diagnostic accuracy from clinical utility and patient benefit
- Design external validation, human interaction, monitoring, and governance for medical imaging AI
A clinical image is evidence inside a care pathway
A chest radiograph, a retinal photograph, a CT series, an ultrasound clip and a pathology slide differ in physics, dimensionality, acquisition and interpretation. The same architecture label tells you little about clinical validity.
Start with the intended use: patient population, user, decision, timing, comparator, consequence, and fallback. That is not a stylistic preference. In January 2018 the FDA authorised the first autonomous AI diagnostic system, through De Novo request DEN180001. It did not authorise a model. It created a device type — "retinal diagnostic software device", class II, product code PIB, now 21 CFR 886.1100. The population, the operating point and the conditions of use went into the classification itself. The rest of this lesson follows that chain, link by link, using studies that have already walked it.
Clinical validity belongs to a defined use context, not to an image model in isolation.
Comparison
Medical imaging tasks support different decisions
Each task requires labels and metrics that match its clinical role. The evidence that settles one row does not settle another. A diagnosis claim rests on a reference standard and a prevalence. A triage claim rests on what happens to the queue behind it. A measurement claim rests on repeatability. A prognosis claim rests on a temporal split, and on the treatments that changed the outcome being predicted. A single reported AUC is compatible with all four and decides none of them.
Detection or triage
Flag images or regions for faster review.
- Sensitivity often prioritized
- Queue effects matter
- Reader remains responsible
- Example: urgent finding alert
Diagnosis or classification
Estimate a condition or category from imaging evidence.
- Reference standard is critical
- Prevalence affects utility
- Differential diagnosis matters
- Example: retinal disease
Measurement or segmentation
Quantify anatomy, lesion, volume, or change.
- Boundary reliability matters
- Repeatability matters
- Editing workflow matters
- Example: tumor volume
Prediction or prognosis
Estimate future outcome using current imaging and context.
- Temporal split essential
- Treatment affects labels
- Causal interpretation risky
- Example: progression risk
Visual
The clinical evidence chain has multiple sources of uncertainty
A reliable study traces uncertainty from acquisition to outcome. Each link now has a named published instrument rather than a general intention.
Step 4 went unnamed the longest. DECIDE-AI is the reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence. It appeared on 18 May 2022, published simultaneously in Nature Medicine and The BMJ. Behind it sat a two-round modified Delphi drawn from 20 pre-defined stakeholder categories: 123 experts in the first round, 138 in the second, 16 at the consensus meeting and 16 in qualitative evaluation. Vasey and colleagues state the result plainly: “The DECIDE-AI reporting guideline comprises 17 AI-specific reporting items (made of 28 subitems) and ten generic reporting items, with an E&E paragraph provided for each.”
It occupies exactly the gap between step 3 and step 5 — between a standalone accuracy figure and the randomised trials that CONSORT-AI governs. A chain that can name a guideline at each link is a chain a reader can audit.
1. Acquire the study
Scanner, protocol, operator, dose, reconstruction, and patient condition shape the image.
2. Define the reference
Pathology, follow-up, consensus reading, or another imperfect standard supplies labels.
3. Evaluate the model
Measure discrimination, calibration, localization, and failure slices.
4. Evaluate interaction
Test how clinicians use, ignore, overrule, or anchor on the output.
5. Evaluate outcomes
Measure time, errors, resource use, equity, and patient consequences.
Key idea
Healthcare datasets contain powerful proxies for the answer
Portable scanner markers, acquisition protocol, site, order type, report language, and treatment artifacts can correlate with disease. Patient images from the same episode can also leak across splits.
Split by patient and time, inspect metadata, and test external sites; a model may predict care processes rather than pathology.
An entire literature once failed this test at once, in public, under maximum pressure. Every machine learning model published or preprinted between 1 January and 3 October 2020 for diagnosing or prognosticating COVID-19 from chest radiographs or CT was reviewed systematically. Roberts and colleagues identified 2,212 studies; 415 survived initial screening; 62 were reviewed in detail. Their verdict, in Nature Machine Intelligence, is one sentence long: “Our review finds that none of the models identified are of potential clinical use due to methodological flaws and/or underlying biases.” None of the 62. One training set used paediatric images as the non-COVID class, so the model had only to separate children from adults. That shortcut scores well on every held-out split drawn from the same two sources.
Clinical shortcuts can be accurate in the dataset and unsafe in deployment.
Case
A pneumonia model that had learned which hospital took the X-ray
Three hospitals, 158,323 chest radiographs, one shortcut. Zech and colleagues measured what a hospital-site proxy does to a pneumonia model. Writing in PLOS Medicine in November 2018, they report that “the prevalence of pneumonia was high enough at MSH (34.2%) relative to NIH and IU (1.2% and 1.0%) that merely sorting by hospital system achieved an AUC of 0.861 (95% CI 0.855-0.866) on the joint MSH-NIH dataset”. No model is involved in that number. It is what the site label alone is worth, before anyone looks at a lung.
The images carry the site too. The same paper reports that “CNNs were able to directly detect hospital system of a radiograph for 99.95% NIH (22,050/22,062) and 99.98% MSH (8,386/8,388) radiographs”. Twelve NIH films and two MSH films, out of thirty thousand, were the network's only difficulty.
A model that knows where a study was taken has a shortcut to who is sick there. No held-out split drawn from those same three sites can expose it. The shortcut is equally valid on both sides of the split.
Figure
Analogy
A laboratory test embedded in a hospital
A new blood test performs well in one laboratory. Adopting it still requires specimen handling, reference ranges, clinician interpretation, quality control, and monitoring across populations.
A blood test has no spatial evidence to read, and reader interaction differs from laboratory reporting. The laboratory still shows why an isolated score is not a clinical system.
Clinical deployment requires analytical, clinical, and workflow validation.
Example
Reference standards are often imperfect and selective
Label provenance changes what a model learns. The largest public chest radiograph corpora are labelled from documentation rather than from the pixels. Johnson and colleagues describe the scale in Scientific Data: “Here we describe MIMIC-CXR, a large dataset of 227,835 imaging studies for 65,379 patients presenting to the Beth Israel Deaconess Medical Center Emergency Department between 2011–2016.” That is 377,110 images, de-identified to satisfy HIPAA Safe Harbor. Each study is paired with the free-text report a practising radiologist wrote during routine clinical care. The report was written to communicate with a treating clinician, not to supervise a classifier.
Size is not accuracy. A board-certified radiologist visually reviewed roughly 700 images from two of the field's standard public datasets, ChestXray14 (112,120 frontal chest films) and MURA (40,561 upper limb radiographs). Luke Oakden-Rayner, who commissioned that review, reports the finding in Academic Radiology: “The ChestXray14 labels did not accurately reflect the visual content of the images, with positive predictive values mostly between 10% and 30% lower than the values presented in the original documentation.” He records hidden stratification and label disambiguation failure alongside it. The MURA normal/abnormal labels reached only 60% sensitivity and 82% specificity for degenerative joint disease. A benchmark can be enormous, public, widely cited, and still mislabelled by twenty points.
- Radiology report: available at scale — 227,835 studies in MIMIC-CXR alone — but it may hedge, omit, or copy text forward. The report-mined labels on ChestXray14 sat 10 to 30 percentage points of positive predictive value below their own documentation.
- Consensus reading: reduces individual variation but remains observer-dependent, and costs a dedicated grading pipeline. IDx-DR's pivotal trial reported 96.1% imageability, so roughly one participant in twenty-five produced no usable image at all.
- Pathology: strong for sampled tissue yet may not represent the entire lesion.
- Follow-up outcome: clinically meaningful but delayed and influenced by treatment.
- Procedure result: available only for patients selected for intervention.
- Administrative code: convenient but shaped by billing and documentation practices. It is the same failure mode as report mining, where the cheap label records a process rather than the pixels.
Steps
Build evidence beyond one retrospective test set
Validate in stages, from technical performance to real clinical use. For one imaging device type these stages are not advice; they are law.
The special controls attached to the retinal diagnostic software device type are codified at 21 CFR 886.1100(b). They require clinical performance testing of sensitivity, specificity, positive predictive value and negative predictive value under anticipated conditions of use. They require confidence intervals on every metric. They require evaluation of variability due to both the user and the image acquisition device, and human factors validation testing of the training programme. They require a protocol defining what level of change to technical specifications could significantly affect safety or effectiveness. Step 2's patient-level split is in the regulation verbatim: “Where multiple samples from the same patient are used, statistical analysis must not assume statistical independence without adequate justification.” That is 21 CFR 886.1100(b)(2)(iii)(A). A leakage lecture with a section number.
Step 5 has a worked example. MASAI was the first randomised controlled trial of AI-supported mammography screening. It randomised 80,033 Swedish women aged 40-80 between April 12, 2021 and July 28, 2022, sponsored by Region Skåne and registered as NCT04838756. AI-supported screening of 39,996 participants produced 244 screen-detected cancers — 6.1 per 1000, 95% CI 5.4-6.9 — from 46,345 screen readings. The 40,024 controls, under standard double reading, produced 203 cancers, 5.1 per 1000 (95% CI 4.4-5.8), from 83,231 readings. The false positive rate was 1.5% in both groups. Lång and colleagues report in The Lancet Oncology that “The screen-reading workload was reduced by 44·3% using AI.”
Note what that sentence pairs. More cancers found, no more false alarms, and half the reading. Three quantities from one prospective randomised design is the shape of a step-5 answer. A discrimination metric from a retrospective test set is not.
1. Lock intended use
Specify modality, population, setting, user, decision, and exclusions.
2. Validate internally
Use patient-level splits, temporal evaluation, calibration, and slice analysis.
3. Validate externally
Test independent sites, devices, protocols, and prevalence.
4. Study readers and workflow
Measure unaided, AI-aided, order effects, disagreement, and workload.
5. Monitor prospectively
Track data quality, drift, overrides, delays, safety events, and outcome proxies.
Key idea
Human plus AI can be worse than either component
Automation bias can make readers accept wrong suggestions. Distrust can make them ignore useful alerts. Interface timing, explanation, alert frequency, and accountability alter combined performance.
Evaluate the actual human-system workflow. Do not infer combined safety by adding model accuracy to clinician accuracy.
The first failure has been measured directly. Twenty-seven radiologists read 50 mammograms carrying BI-RADS suggestions from a purported AI system, which was incorrect for 12 of the 40 test cases. When the suggestion was wrong, the proportion of correctly categorised mammograms fell from 79.7% to 19.8% among inexperienced readers, from 81.3% to 24.8% among moderately experienced readers, and from 82.3% to 45.5% among the very experienced. Dratsch and colleagues state the conclusion in Radiology: “The results show that inexperienced, moderately experienced, and very experienced radiologists reading mammograms are prone to automation bias when being supported by an AI-based system.” Experience blunted the collapse. It did not prevent it. A reader who was right four times in five became wrong more often than right.
The comparison a vendor quotes is rarer than it looks. Liu and colleagues report that “our search identified 31 587 studies, of which 82 (describing 147 patient cohorts) were included”, and that “an out-of-sample external validation was done in 25 studies, of which 14 made the comparison between deep learning models and health-care professionals in the same sample”. Tschandl and colleagues then ran the interaction itself, for skin cancer. They found that “good quality AI-based support of clinical decision-making improves diagnostic accuracy over that of either AI or physicians alone, and that the least experienced clinicians gain the most from AI-based support” — and also that “faulty AI can mislead the entire spectrum of clinicians, including experts”. CONSORT-AI, published in 2020, “includes 14 new items that were considered sufficiently important for AI interventions that they should be routinely reported in addition to the core CONSORT 2010 items”. Among them are the human-AI interaction and an analysis of error cases.
Clinical benefit is an interaction effect, not a property of the model alone.
Example
Practice: plan an external validation for retinal screening
A model trained on one camera network will be introduced in clinics with new devices, referral patterns, and disease prevalence. You do not have to invent the plan. One has already been written, submitted and accepted, and it is public.
IDx-DR became the first autonomous AI diagnostic system authorised by the FDA, through De Novo request DEN180001, dated January 12, 2018. Its pivotal trial enrolled 900 participants at 10 primary care sites, of whom 819 were fully analysable. Prevalence of more-than-mild diabetic retinopathy was 23.8% (198/819), and 96.1% of participants were imageable at all. The agency's own decision summary records the result under "Summary of Benefits": “The pivotal clinical study, which enrolled a total of 900 participants, demonstrated observed sensitivity for mtmDR 87.4%, with observed specificity of 89.5%.” Enrichment-corrected, sensitivity was 87.2% (95% CI 81.8-91.2) and specificity 90.7% (95% CI 88.3-92.7); PPV was 72.7% (173/238) and NPV 95.7% (556/581). Abràmoff and colleagues reported the same trial, registered as NCT02963441, in npj Digital Medicine.
Write your plan against that document, and for each line ask what it cost to produce.
- Define patient-level inclusion, exclusion, and reference-standard rules: the De Novo summary fixes all three, including which reading centre graded the images and by what protocol.
- Select site, device, age, pigmentation, and image-quality slices; the trial ran at 10 primary care sites and reported observed and enrichment-corrected operating points separately rather than blending them.
- Report sensitivity, specificity, calibration, ungradable rate, and referral workload — here 87.4% and 89.5% observed, 96.1% imageability, and a PPV of 72.7% (173/238) at a prevalence of 23.8%.
- Design a reader workflow for uncertain or poor-quality images: 819 of the 900 enrolled participants were fully analysable, and that gap is where the workflow becomes visible in the numbers.
- Specify monitoring triggers and who can suspend the system; 21 CFR 886.1100(b) already demands a protocol defining what level of change to technical specifications could significantly affect safety or effectiveness.
Position
Matching the radiologist is not the comparison that decides anything
The headline for medical imaging AI is a duel. The model scored this, the radiologists scored that, and one of them won. Two things are wrong with the framing.
The first is how rare the duel is. Liu and colleagues screened 31,587 studies and included 82, of which 25 carried an out-of-sample external validation. Only 14 made the comparison against health-care professionals in the same sample. In those 14, the pooled sensitivity was “87·0% (95% CI 83·0-90·2) for deep learning models and 86·4% (79·9-91·0) for health-care professionals”. Six tenths of a point, from fourteen studies, with overlapping intervals, is not a verdict on a profession.
The second is that the duel is not the deployment. The deployments in question put the model beside a reader, not in place of one, and the sign of the effect is set there. Tschandl and colleagues found that good support raised accuracy above either party alone, and that the least experienced clinicians gained most — and also that “faulty AI can mislead the entire spectrum of clinicians, including experts”. Dratsch and colleagues put a size on the downside. Correct categorisation of mammograms fell from 82.3% to 45.5% among very experienced readers when the AI suggestion was wrong, and from 79.7% to 19.8% among inexperienced ones. The same standalone score can precede a gain or a harm, depending on how the output is presented and who is reading it.
There is a third problem underneath both. The standalone score may not be measuring disease at all. Zech and colleagues found pneumonia prevalence of 34.2% at MSH against 1.2% and 1.0% at NIH and IU. Sorting images by hospital system alone reached an AUC of 0.861 (95% CI 0.855-0.866), and the networks recovered the source hospital from the image itself for 99.95% and 99.98% of radiographs. Roberts and colleagues reviewed 62 COVID-19 imaging models in detail and found none of potential clinical use.
None of this says the models are weak. The standalone comparison is one link in a five-stage chain: acquisition and the reference standard come before it, interaction and patient outcomes after. DECIDE-AI's 17 AI-specific items govern the early live evaluation. CONSORT-AI's 14 added items govern the trial, including the human-AI interaction and an analysis of error cases. MASAI shows what the end of the chain reads like — 6.1 versus 5.1 cancers per 1000, an identical 1.5% false positive rate, and a screen-reading workload cut by 44·3%. When a vendor quotes a comparison against clinicians, ask which sample, against whom, and what happened when a clinician was in the room.
Only 14 of the 82 reviewed studies compared a model with clinicians on the same sample, and a comparison is not an interaction.
Key takeaways
- Medical imaging AI must be defined by modality, population, user, decision, timing, and clinical consequence. DEN180001 authorised a device type and an operating point under 21 CFR 886.1100, not a model.
- Acquisition artifacts, site metadata, care pathways, and patient overlap create dangerous shortcuts. Sorting Zech's radiographs by hospital system alone reached an AUC of 0.861, and 62 reviewed COVID-19 imaging models yielded none of potential clinical use.
- Reference standards differ in validity, availability, selectivity, and uncertainty. ChestXray14's report-mined labels ran 10 to 30 percentage points of positive predictive value below their own documentation, and MURA's reached 60% sensitivity for degenerative joint disease.
- External and temporal validation test transfer beyond the development setting. 21 CFR 886.1100(b)(2)(iii)(A) makes patient-level statistical dependence and per-metric confidence intervals a legal requirement rather than a preference.
- Human-AI interaction can improve or degrade performance and must be measured directly. Correct categorisation fell from 82.3% to 45.5% among very experienced readers, and from 79.7% to 19.8% among inexperienced ones, when the AI suggestion was wrong.
- Clinical deployment requires monitoring, governance, fallback, incident response, and evidence tied to patient outcomes: DECIDE-AI for the early live stage, CONSORT-AI for the trial, and MASAI's 6.1 versus 5.1 cancers per 1000 at an unchanged 1.5% false positive rate for the outcome.