Industry Transformation
The AI Healthcare Revolution: How Machine Intelligence Is Rebuilding Medicine From the Ground Up
Ambient scribes in millions of visits, 1,450 AI devices on the FDA's list, and the first AI-designed drug dosed in Phase III in September 2026 — where medical AI is working, where it failed, and why validation is the whole story.

Gabriele Masetti ·
The paperwork of medicine, quietly rewritten
Walk into an exam room in 2026 and the most visible change isn't a robot or a screen full of neural-network diagrams. It's a phone propped on the desk, listening. Ambient AI scribes — Abridge, Microsoft's Dragon Copilot (the product formerly known as Nuance's DAX), and Epic's own "Art for Clinicians" — now sit in on millions of patient visits, transcribing the conversation and drafting a clinical note before the patient has left the parking lot. This is the least glamorous corner of medical AI and, arguably, the one already doing the most good at scale.
The evidence isn't a press release. Kaiser Permanente's Division of Research, publishing in NEJM Catalyst and in follow-up work through Permanente Medicine, tracked The Permanente Medical Group's rollout of ambient scribes starting in late 2023: within a year, more than 2.5 million patient encounters had used the tool, and physicians collectively saved close to 16,000 hours of documentation time, with the biggest gains concentrated in mental health, emergency medicine, and primary care — exactly the specialties that report the worst burnout. Physician and patient surveys were positive, and Kaiser has continued publishing quality-assurance data as adoption scaled, rather than declaring victory after a pilot.
The commercial market has moved just as fast. Abridge, which built its own foundation models rather than wrapping a general-purpose LLM, has been named Best in KLAS for ambient AI in both 2025 and 2026 and raised a $300 million Series E in June 2025 at a $5.3 billion valuation, then added a $316 million extension in April 2026 at the same valuation — serious money chasing a genuinely unsexy problem: the note.
Microsoft, which bought Nuance for $19.7 billion in 2022, folded DAX into Dragon Copilot and has since widened it well past the note. At HIMSS in March 2026 the company repositioned Dragon Copilot as a clinical assistant rather than a scribe, extended it to nurses and radiologists, opened it to third-party agents, and put its installed base at more than 100,000 clinicians. Microsoft's own product page now lists physician availability in ten countries, nurse availability in the United States only, and the radiology experience still in preview. Epic, not wanting to cede the point of integration, entered the market itself in August 2025 with a native tool built on Microsoft's Dragon Ambient AI and trained against Cosmos, Epic's aggregated database of roughly 300 million patient records.
Three well-capitalized players racing to own the same five minutes of a doctor's day is a reasonable signal that documentation burden — not diagnosis — was medicine's most tractable AI problem, because the ground truth (what was said in the room) is unambiguous and the failure mode (a wrong word in a note) is correctable by a human before it's signed.
Where imaging AI has actually earned its clearances
Diagnostic imaging is where AI in medicine has the longest track record and the clearest regulatory paper trail. The FDA's public list of AI-enabled medical devices held roughly 1,450 entries at the end of 2025 — the running total comes from third-party snapshots, since the agency does not publish one — with about 295 of them cleared during 2025 alone. Radiology accounts for roughly 76% of the list, about 1,100 devices, because imaging is data-rich, well-labeled, and structurally suited to pattern recognition.
The landmark case is IDx-DR, now rebranded LumineticsCore. In 2018 it became the first FDA De Novo authorization for a fully autonomous diagnostic AI — no physician needs to interpret the image at all. A primary-care technician photographs a patient's retina with a Topcon camera, and the software itself renders a diagnosis of more-than-mild diabetic retinopathy, no ophthalmologist required for that determination.
Its pivotal trial hit 87% sensitivity, 90% specificity, and a 96% imageability rate, and the system is now deployed at more than 20 U.S. health systems. It matters not because the algorithm is exotic but because the regulatory category itself — a machine allowed to make an unsupervised diagnostic call — didn't previously exist.

Pathology got its own milestone in 2021, when Paige Prostate became the first FDA-authorized AI product in digital pathology, cleared to help detect cancer in prostate needle biopsies. In validation, pathologists using the tool cut false negatives by roughly 70% and false positives by about 24% compared with unassisted reads, and — the detail that should reassure skeptics of AI-as-deskilling — generalist pathologists using the software performed comparably to unassisted subspecialists. That's the pattern imaging AI keeps demonstrating: it narrows the gap between an average clinician and a top specialist more reliably than it improves on the best human reader alone.
| Tool | Metric | Improvement |
|---|---|---|
| Paige Prostate | False negatives | cut by ~70% |
| Paige Prostate | False positives | cut by ~24% |
| DeepMind mammography (US) | False positives | reduced 5.7% |
| DeepMind mammography (US) | False negatives | reduced 9.4% |
| DeepMind mammography (UK) | False positives | reduced 1.2% |
| DeepMind mammography (UK) | False negatives | reduced 2.7% |
Google DeepMind's mammography work, published in Nature and built with Cancer Research UK's Imperial Centre, Northwestern University, and Royal Surrey County Hospital, trained on more than 90,000 de-identified mammograms across the UK and US and, in retrospective reader studies, reduced false positives by 5.7% (US) and 1.2% (UK), and false negatives by 9.4% (US) and 2.7% (UK), against a panel of human experts.
DeepMind was upfront about the catch: the study used images predominantly from one manufacturer's equipment, and prospective clinical trials — not retrospective reader studies — are what would actually establish real-world performance. That caveat has proven durable; screening-AI deployment has moved more slowly than the 2020 headlines implied, precisely because retrospective accuracy and prospective clinical benefit are different questions.
AlphaFold and the slow grind from prediction to pill
The 2024 Nobel Prize in Chemistry went to Demis Hassabis and John Jumper of Google DeepMind for AlphaFold, alongside David Baker for computational protein design — the first time a Nobel has been awarded for an achievement built on AI. The freely released AlphaFold Protein Structure Database has since been accessed by more than 3 million researchers in over 190 countries, including over a million users in low- and middle-income countries, and roughly 30% of the research citing it is disease-focused.
That reach is real and it is not hype: predicting a protein's 3D structure from its amino acid sequence, a problem biologists worked on for decades, is now something a model does in minutes.
What AlphaFold has not yet done is put a new drug on a pharmacy shelf. Isomorphic Labs, DeepMind's drug-discovery spinout, raised $600 million in March 2025, has partnerships with Eli Lilly and Novartis, and in February 2026 released IsoDDE, a unified design engine combining structure prediction with binding-affinity modeling — but it has still not given a molecule to a person. Speaking at Davos in January 2026, Demis Hassabis said the first clinical trials would come by the end of 2026, a year later than he had previously forecast, and that target was still ahead of the company in September 2026.
Insilico Medicine's molecule for the same disease has gone further than anything else in the field. INS018_055, now named rentosertib, reported Phase IIa results in Nature Medicine on 3 June 2025: 71 patients with idiopathic pulmonary fibrosis across 22 Chinese sites, twelve weeks of dosing, and a mean gain of 98.4 mL in forced vital capacity on 60 mg once daily against a 20.3 mL decline on placebo. On 10 September 2026 Insilico dosed the first patient in GENESIS-IPF-3, a 52-week, randomized, placebo-controlled Phase III aiming at 320 patients across 47 centers in China — the first time a compound with both an AI-selected target and an AI-generated structure has reached that stage.
Both are meaningful firsts. Neither is proof yet that AI-designed drugs work better, cheaper, or faster than the conventional pipeline — a Phase III that has just enrolled its first patient is a starting gun, not a result. The honest read on AlphaFold's impact is: transformative for structural biology as a research tool, unproven — so far — as a drug-approval accelerant. Those are different claims, and the industry's own press cycle tends to blur them.
The failures that should discipline the optimism
The cautionary case that every health system now cites by name is the Epic Sepsis Model, a proprietary early-warning algorithm embedded in the Epic EHR and adopted at hundreds of U.S. hospitals without independent external validation. When outside researchers finally got access to test it, the results were bad: a University of Michigan validation found the model missed two-thirds of sepsis cases it should have caught, and a separate 6-hour-window analysis found sensitivity of just 14.7% with a positive predictive value of 7.6% — meaning the overwhelming majority of its alerts were false alarms, while it silently missed most real cases.
The damage wasn't that the model was wrong; it's that a widely deployed clinical tool ran for years on vendor-reported performance before anyone checked its real-world calibration. That is now the central argument for mandatory external validation of any clinical prediction model before wide deployment, not just at FDA clearance but continuously, since a model trained on one hospital's data drifts when it hits another hospital's patient mix, staffing, and documentation habits.
| System | What went wrong | Documented outcome |
|---|---|---|
| Epic Sepsis Model | Deployed without independent validation | Missed two-thirds of sepsis cases; 14.7% sensitivity, 7.6% PPV |
| IBM Watson for Oncology | Trained mainly on one hospital's protocols | MD Anderson spent ~$62 million before shelving in 2016 |
| Commercial risk-prediction algorithm | Used spending as a proxy for illness severity | Correcting bias raised flagged Black patients from 17.7% to 46.5% |
IBM Watson for Oncology is the older, more expensive version of the same lesson. The University of Texas MD Anderson Cancer Center spent approximately $62 million on an Oncology Expert Advisor project before shelving it in 2016, and internal IBM documents later surfaced by STAT News in 2018 showed the system had at times recommended "unsafe and incorrect" cancer treatments in testing.
Part of the failure was structural: Watson had been trained largely on cases and protocols from Memorial Sloan Kettering, and its recommendations didn't transfer cleanly to other institutions' patient populations and treatment norms — a training-data generalization problem, not a one-off bug. No patients were harmed, because the flawed outputs were caught before clinical use, but the episode remains the reference point for what happens when an AI system's marketing outpaces its validation.
Bias is the third recurring failure mode, and the clearest documented case is Ziad Obermeyer's 2019 Science paper dissecting a widely used commercial risk-prediction algorithm that hospitals used to flag patients for extra care management. The algorithm used historical health-care spending as a proxy for illness severity — a reasonable-sounding shortcut that broke down because, at equal levels of actual sickness, less money was historically spent on Black patients due to unequal access to care.
The result: at any given risk score, Black patients were significantly sicker than white patients, and the algorithm was referring far fewer Black patients than it should have for extra care — the study found correcting the bias would have raised the share of Black patients flagged for extra help from 17.7% to 46.5%. Worked with the manufacturer, Obermeyer's team showed that retraining on a more direct measure of health need cut the racial bias by 84%.
The fix was tractable once the flaw was found — but the flaw had been running in production, invisibly, for an unknown period before an academic audit caught it. That's the pattern across all three failures: the tools weren't obviously broken until someone outside the vendor tested them against reality.
Regulation is trying to catch up, not keep up
The FDA's approach has shifted from clearing static algorithms one at a time toward managing AI as products that keep changing after clearance. Its January 2025 guidance on lifecycle management for AI-enabled device software functions, still marked "Draft — Not for implementation" in September 2026, sets out a total product life cycle framework, asking manufacturers for model description, data lineage, documented performance tied to specific claims, and bias analysis. The companion document is already final: the FDA's guidance on predetermined change control plans for AI-enabled device software functions, issued in August 2025, lets a manufacturer pre-agree exactly how and when a model may update itself without a new marketing submission.
The agency has also said it will explore ways to identify and tag entries on that list which are built on foundation models, from large language models to multimodal architectures, and asks sponsors to put the information in their public summaries — an acknowledgment that generative AI doesn't fit neatly into a framework designed around fixed classifiers analyzing images, and a tag that had not yet appeared on the list by September 2026. None of this resolves the harder open question: how do you validate a system whose behavior can legitimately change every few months, deployed across health systems with different patient populations, staffing models, and data pipelines than whatever validation set the manufacturer used?
The throughline
Set the record straight against the two competing narratives. It is not true that AI is quietly failing in medicine — the diabetic retinopathy screener, the prostate pathology tool, and the ambient scribes are running at real scale, with published outcomes, saving real clinician hours and catching real disease. It is also not true that medicine is being "rebuilt from the ground up" by intelligent machines, the framing this piece's own headline promises.
What's actually happening is narrower and more useful: AI is winning in domains with clean ground truth and forgiving failure modes — a retinal image either shows the lesion or it doesn't, a note either matches the recorded conversation or gets corrected before signing — and it is failing, sometimes expensively, wherever a model's era of training data quietly stops matching the population it's now judging, from Epic's sepsis alerts to Watson's treatment plans to a cost-based risk score that mistook unequal access for good health.
AlphaFold's Nobel is real; so is the fact that no AI-designed drug has yet cleared a Phase III trial to prove the pipeline works end to end, and the first one to get there only started dosing in September 2026. The honest 2026 assessment isn't revolution or hype — it's that the technology is exactly as good as the validation discipline applied to it, and that discipline has so far been wildly uneven between a well-audited retina scanner and a sepsis model that ran for years before anyone checked it against outside data.