Skip to content
AI.info

Future Horizons

AI and Drug Discovery: How Machine Intelligence Is Compressing Decades of Medical Research Into Years

Rentosertib opened Phase III in July 2026 and two more AI-designed molecules are on the FDA's review calendar. AI drug discovery has moved past compressed timelines to filings, approvals and the cost of carrying a pipeline.

AI and Drug Discovery: How Machine Intelligence Is Compressing Decades of Medical Research Into Years

Gabriele Masetti ·

The Arithmetic of Drug Discovery

Every drug that reaches patients represents a statistical miracle. The industry average for bringing a new molecular entity from initial discovery to regulatory approval is 12 to 15 years. The average cost is $2.6 billion, including the cost of failures. And the failure rate is staggering: roughly 90% of drug candidates that enter Phase I clinical trials never reach patients, most of them eliminated not by toxicity but by efficacy — they simply do not work well enough.

Behind every approved drug is a graveyard of thousands of molecules that were synthesised, tested, found wanting, and discarded. The process of drug discovery is, in its traditional form, a systematic exploration of an almost incomprehensibly large chemical space — the estimated number of drug-like molecules that could theoretically be synthesised runs to 10 to the power of 60 — using biological tools that are powerful but slow, expensive, and fundamentally serial.

Artificial intelligence is attacking the arithmetic of drug discovery on every front simultaneously. It is compressing the time from target identification to clinical candidate selection from years to months. It is reducing the number of compounds that need to be synthesised and tested in the laboratory by predicting their properties computationally. It is identifying new disease targets that would never have been found by traditional methods. And it is improving the prediction of clinical success, the most stubborn source of the industry's enormous failure rate.

The most important milestone in this story came in 2023, when Insilico Medicine received approval from Chinese regulators to advance ISM001-055 — since renamed rentosertib, the first AI-discovered drug candidate — into Phase II clinical trials for idiopathic pulmonary fibrosis.

Its Phase IIa results appeared in Nature Medicine on 3 June 2025, and on 7 July 2026 Insilico opened a 320-patient Phase III trial. By September 2026 two further AI-designed molecules were sitting in front of the FDA with accepted marketing applications. The era of AI drug discovery had moved from promise to regulatory filings.

The Biology of Drug Discovery: Why the Problem Is Hard

To understand why AI matters so much to drug discovery, it helps to understand why drug discovery is so hard.

A drug works by interacting with a biological target — typically a protein that plays a role in a disease process.

Finding or designing a small molecule that binds to the right protein, in the right way, with sufficient potency, without binding to the wrong proteins (causing side effects), without being metabolised too quickly or too slowly, without being toxic to the liver or kidneys, and with the physical properties needed to be absorbed from the gut and reach the tissue where the disease is occurring.

That is the challenge that the industry has been working on for a century, with tools that are incrementally better but fundamentally limited in their ability to predict what a molecule will do before it is made.

The protein folding problem — predicting the three-dimensional structure of a protein from its amino acid sequence — was one of the central unsolved problems of structural biology for fifty years. Without knowing a protein's structure, it is enormously difficult to design molecules that interact with it in specific ways.

AlphaFold, DeepMind's protein structure prediction system, solved this problem with stunning accuracy in 2020 and released its predictions for virtually every protein in the human proteome in 2021. It is hard to overstate the significance of this achievement for drug discovery: it compressed the timeline for structure-based drug design from years to weeks for any given target.

AlphaFold was not the first AI application to drug discovery, but it was the one that demonstrated, beyond reasonable doubt, that AI could solve hard biological problems that traditional methods could not. It opened the floodgates for investment and research across the entire drug discovery pipeline.

Generative AI and the Chemical Space Problem

The challenge of finding a molecule that hits a specific target with the required properties is, in abstract mathematical terms, a search problem. The space being searched — the universe of possible drug-like molecules — is vast beyond conventional comprehension.

The traditional approach to this search was to build libraries of known compounds, screen them against targets of interest, and iteratively modify the most promising hits. The approach is powerful but inefficient: even the largest compound libraries contain only a tiny fraction of possible drug-like molecules, and the most promising regions of chemical space may lie far from any compound that has ever been synthesised.

Generative AI offers a fundamentally different approach. Instead of searching through a fixed library, generative models — typically variants of the diffusion models used in image generation or of transformer architectures adapted for molecular graphs — can design entirely new molecules optimised for specific properties. Given a target protein and a set of desired properties (potency, selectivity, oral bioavailability, metabolic stability, synthetic accessibility), a generative model can propose novel molecular structures that have never been synthesised before, optimised across multiple objectives simultaneously.

Novartis provided one of the most striking demonstrations of this capability. Using generative AI, researchers computationally designed 15 million potential compounds optimised for a particular target, then applied predictive models to assess brain penetration and other properties, narrowing the candidate pool to a set of roughly 60 molecules for laboratory synthesis and testing. Traditional approaches would have required synthesising thousands of molecules to achieve a comparable degree of optimisation. The AI-assisted approach delivered a potent, brain-penetrant molecular scaffold for experimental validation.

The efficiency gain is not just about the number of compounds. It is about the quality of the exploration. AI-guided design can explore regions of chemical space that are far from known compounds and traditional medicinal chemistry intuition — regions that might harbour drug candidates that human chemists would never have considered.

The Clinical Era: First-Generation AI Drugs in Trials

The ultimate test of AI drug discovery is clinical: do AI-designed drugs work in humans? The answer, based on early evidence, is cautiously encouraging.

Insilico Medicine's rentosertib, formerly ISM001-055, for idiopathic pulmonary fibrosis achieved the entire journey from target identification to clinical candidate using AI in 18 months, at a cost of approximately $2.6 million — compared to an industry average of 4-5 years and $300-600 million for the same stages using traditional methods.

The Phase IIa trial, published in Nature Medicine on 3 June 2025, randomised 71 patients across three dose arms and placebo over 12 weeks. Patients on 60 mg once daily gained a mean 98.4 mL of forced vital capacity; the placebo group lost a mean 20.3 mL.

On 7 July 2026 Insilico dosed the first patients in a Phase III trial of 320 IPF patients across 47 Chinese centres, running 52 weeks — the first drug discovered by generative AI to reach a pivotal trial.

The Phase IIa results gave us the confidence to advance Rentosertib into larger and longer clinical testing. — Carol Satler, Senior Vice President for Clinical Development, Non-Oncology, Insilico Medicine

Metric Traditional AI-driven (rentosertib)
Time to clinical candidate 4-5 years 18 months
Cost $300-600 million ~$2.6 million

AbSci developed antibodies designed using deep learning that entered clinical trials in 2024. Schrödinger's physics-based AI platform produced zasocitinib (TAK-279), a tyrosine kinase 2 inhibitor that has now cleared pivotal testing. In the LATITUDE PsO 3001 and 3002 trials, 71.4% and 69.2% of patients on zasocitinib reached a static Physician Global Assessment score of clear or almost clear at week 16, against 32.1% and 29.7% on apremilast and 10.7% and 12.6% on placebo.

A separate head-to-head trial, LATITUDE Atlas, reported on 11 June 2026 that more than 35% of patients on zasocitinib 30 mg reached complete skin clearance at week 16, more than 2.5 times the rate on deucravacitinib 6 mg.

The FDA accepted Takeda's new drug application under Priority Review on 14 September 2026, with a target action date in the first quarter of 2027.

Relay Therapeutics' RLY-4008, an FGFR2 inhibitor designed using AI-assisted conformational analysis, is now lirafugratinib, licensed to Elevar Therapeutics in December 2024 for $5 million upfront and up to $495 million in milestones.

In the pivotal ReFocus cohort it produced an independently assessed objective response rate of 46.5% and a median duration of response of 11.8 months. The FDA accepted Elevar's application under Priority Review on 30 March 2026 and set a target action date of 27 September 2026 — one of the first approval decisions on a molecule designed with computational and AI methods.

Candidate Company Indication/type Stage, September 2026
Rentosertib (ISM001-055) Insilico Medicine Idiopathic pulmonary fibrosis Phase III opened 7 July 2026, 320 patients
Zasocitinib (TAK-279) Schrödinger / Takeda Tyrosine kinase 2 inhibitor NDA accepted under Priority Review, 14 September 2026
Lirafugratinib (RLY-4008) Relay Therapeutics / Elevar FGFR2 inhibitor FDA target action date 27 September 2026
AbSci antibodies AbSci Deep learning-designed antibodies Entered clinical trials (2024)

Counting AI-designed candidates in the clinic is harder than the headline figures suggest: published tallies run from a few dozen to well over a hundred, because no two agree on whether a molecule qualifies when a model generated its structure, merely ranked it, or only picked its target.

What is countable is what reached regulators — between July and September 2026, three AI-designed molecules moved into a pivotal trial or onto an FDA review calendar. The leading AI drug biotechs — Iambic, Generate Biomedicines, Recursion — assembled clinical portfolios in a handful of years rather than a decade.

The question of whether these candidates would show the clinical success rates AI developers claimed is now being answered in instalments rather than all at once. Zasocitinib met its endpoints and beat two active comparators; lirafugratinib reached the FDA on a 46.5% response rate; rentosertib has only just begun its pivotal trial.

Three readouts are not a success rate, and the industry average Phase III success rate is approximately 50% — meaning that of drugs entering the final pivotal trial stage, about half ultimately receive approval. If AI-designed drugs achieve substantially higher Phase III success rates, it will confirm the most optimistic claims for the technology. If they fail at rates similar to traditional drugs, the value proposition shifts from reducing failures to reducing the cost and time of getting to the same failure rates.

Self-Driving Laboratories: Closing the Loop

AI drug discovery is not just about computational design. It is also about closing the loop between design and experimental testing — and doing so in a way that is faster, cheaper, and more reproducible than traditional laboratory work.

Self-driving laboratories, also called autonomous laboratories or robotic laboratories, integrate AI-guided experimental design with automated laboratory equipment that can execute experiments, record results, and feed those results back into the AI system for the next iteration. The design-make-test-learn cycle that is at the heart of drug discovery can, in a self-driving laboratory, operate continuously and autonomously.

The University of Toronto's Acceleration Consortium, a major research programme in autonomous chemistry, demonstrated self-driving laboratories that could run thousands of experiments per day and identify optimal synthesis conditions for novel molecules in hours rather than months. Emerald Cloud Lab in the United States offers cloud-based access to automated laboratory infrastructure that allows computational scientists to run physical experiments remotely through software interfaces.

AstraZeneca, Pfizer, Novartis, and other large pharmaceutical companies are investing heavily in integrated automation platforms that connect computational design tools with robotic synthesis and testing systems. The goal is not just efficiency but consistency: automated laboratories produce data that is more reproducible, better documented, and more easily integrated with machine learning models than data from traditional manual experiments.

The implications extend beyond efficiency. Self-driving laboratories make it possible to explore the design-make-test-learn cycle at a frequency that was previously impossible, iterating through molecular designs on a timescale of days rather than months. That changes the fundamental economics of the early discovery phase and opens possibilities for personalised medicine applications — designing molecules optimised for individual patients' specific genetic profiles and disease characteristics — that would be impractical using traditional methods.

The Target Identification Revolution

Finding the right protein target for a disease — the point where AI-designed molecules should bind to have a therapeutic effect — has historically been one of the most difficult and uncertain steps in drug discovery. Disease biology is complex, incompletely understood, and full of situations where targeting an apparently logical protein produces disappointing results in practice.

AI is transforming target identification in two ways. First, by integrating and analysing the enormous datasets that have accumulated from genomics, proteomics, transcriptomics, and clinical studies to identify patterns that point toward novel disease mechanisms. Second, by enabling in silico experimentation — simulating the effect of perturbing thousands of genes or proteins in digital models of disease cells — that would be impractical in the laboratory.

Microsoft Research's work with the Institute for Neurodegenerative Diseases at UCSF used large quantitative models — a class of AI trained on experimental data from biological simulations — to explore therapeutic targets for Alzheimer's disease and related conditions. By simulating the perturbation of thousands of genes in computational models of disease-relevant cell types, researchers could identify targets with predicted therapeutic relevance before committing to any laboratory experiments.

Genentech and other large biotechs are using graph neural networks trained on protein interaction data to identify proteins that, when targeted, could have cascading beneficial effects on disease-relevant pathways. These network-based approaches can identify targets that would be invisible to traditional single-target thinking but that may offer superior efficacy precisely because they engage with biology's inherent complexity.

The Regulatory Challenge: Teaching the FDA About AI

The accelerating clinical output of AI drug discovery is creating unprecedented challenges for regulatory agencies. The FDA was not designed to evaluate drugs whose entire development journey — from target identification through clinical candidate selection — was guided by algorithms.

In January 2025, the FDA published draft guidance on a risk-based credibility assessment framework for AI models used in regulatory contexts. The guidance emphasises "context of use" — the specific decision the AI model is being used to make — and requires ongoing performance evaluation as the model encounters new data and new situations. That is a marked departure from traditional regulatory frameworks, which evaluate a fixed method or process rather than an adaptive system.

The specific regulatory challenges are substantial. How should the FDA evaluate the evidence base for an AI-designed molecule when the design process itself is not fully transparent — when the generative model that proposed the molecule cannot fully explain why it proposed it? How should regulators evaluate the reliability of AI-predicted properties when the prediction method is novel and lacks the decades of validation data that traditional methods have accumulated?

The EU AI Act adds a further layer of complexity for drug developers operating in European markets, though on a later clock than first legislated. High-risk AI applications in healthcare — which will include many drug discovery applications — face conformity assessment requirements, data governance standards, and human oversight mandates that will require pharmaceutical companies to document their AI systems in ways that current practice does not require.

Under the Digital Omnibus agreed by EU negotiators on 6 May 2026, the obligations for the stand-alone high-risk systems listed in Annex III were pushed to 2 December 2027, and those for AI embedded in regulated products under Annex I to 2 August 2028.

None of these regulatory challenges are insurmountable. They are the normal friction of new technology encountering regulatory frameworks designed for older technology. But they will add cost and time to AI drug development pipelines, and navigating them effectively will require investment in regulatory science capabilities that most AI drug biotechs do not currently have.

The Competitive Landscape: Big Pharma Meets AI Biotech

The entry of AI into drug discovery has created a fascinating competitive dynamic between established pharmaceutical companies and a new generation of AI-native biotechs.

The large pharmaceutical companies have enormous advantages: deep biological expertise accumulated over decades, large proprietary datasets from years of research, established clinical development infrastructure, regulatory relationships, and the financial resources to absorb the many failures that are inherent to drug development. They are now actively acquiring AI capabilities — through partnerships, acquisitions of AI biotechs, and internal investment in data science and computational capabilities.

Bristol-Myers Squibb, AstraZeneca, Roche, and Novartis all have substantial internal AI drug discovery programmes and have signed major partnerships with AI biotechs. Pfizer used AI extensively in its COVID-19 antiviral development and has committed to making AI central to its R&D strategy. Eli Lilly acquired multiple AI drug discovery capabilities and has AI-designed candidates progressing through its pipeline.

The AI-native biotechs — Recursion, Schrödinger, Insilico Medicine, Exscientia (now part of Recursion after a 2024 merger), Generate Biomedicines, Iambic, AbSci — have different advantages: speed, focus, lack of legacy infrastructure to navigate, and in some cases truly novel platform capabilities that incumbents cannot easily replicate. They are building clinical track records that will determine their long-term value, and the next two to three years of clinical readouts will establish who has genuinely transformative technology versus who has impressive computational capabilities that do not translate to clinical success.

None of those advantages is proof against the ordinary arithmetic of biotech. Recursion, which absorbed Exscientia in a 2024 merger and entered 2025 with roughly 800 employees, deprioritised three clinical-stage programmes — in neurofibromatosis type 2, cerebral cavernous malformation and C. difficile infection — then cut about 20% of its staff in June 2025, narrowing its research to oncology and rare disease and pushing its cash runway into the fourth quarter of 2027.

A platform that can generate candidates faster than its owner can afford to run them is a financing problem as much as a scientific advance.

The Most Challenging Diseases: Where AI Falls Short

The enthusiasm for AI drug discovery should not obscure its current limitations, particularly in the disease areas that cause the most human suffering.

Neurodegenerative diseases — Alzheimer's, Parkinson's, ALS, Huntington's — remain the graveyard of drug development. The biology is incompletely understood, the animal models are poor predictors of human disease, and the clinical development timeline is extraordinarily long because outcomes take years to measure. AI's ability to identify targets and design molecules is limited by the quality and relevance of the data it learns from. In neurodegenerative disease, that data is often inadequate.

Mental health is similarly difficult. The lack of clear biomarkers for psychiatric conditions, the heterogeneity of patient populations, and the inadequacy of existing animal models all limit the value that AI can currently add. The tools are being developed — patient stratification based on genomics and digital biomarkers, AI-guided trial design, more sophisticated models of neural circuit function — but the hard biological problems remain unsolved.

Antimicrobial resistance is a different kind of challenge. The science of identifying and designing novel antibiotics is increasingly tractable with AI tools. The challenge is economic: antibiotics that are used sparingly to preserve their effectiveness generate limited revenue for pharmaceutical companies, and the investment required to develop them is not easily recouped. AI can design better antibiotics. It cannot fix the economics of antibiotic development, which is a market failure that requires policy solutions, not better algorithms.

The $25 Billion Market and What It Means

The AI drug discovery market was valued at approximately $5-7 billion in 2025 and is projected to grow to $25 billion or more by the early 2030s. These projections are driven not just by the clinical success of early AI-designed drugs but by the fundamental economics: if AI can reduce the cost and time of drug development by even 30-40% while improving the probability of clinical success, the value creation is enormous.

The implications for patients are the aspect most worth focusing on. The current economics of drug development — enormous cost, long timelines, high failure rates — mean that many diseases are inadequately served by the pharmaceutical industry because the expected return on investment does not justify the development cost. Rare diseases, diseases that disproportionately affect populations in lower-income countries, diseases where the biology is poorly understood — all of these are under-served because the arithmetic doesn't work for traditional drug development.

If AI substantially changes that arithmetic — reducing development costs from hundreds of millions to tens of millions, compressing timelines from decades to years — it opens the possibility of addressing a far wider range of diseases than the current pharmaceutical business model allows. That is a consequential possibility. It is also a fragile one: it depends on the clinical validation that is currently underway, on regulatory frameworks that can evaluate novel development pathways, and on business models that make AI-enabled drug development economically viable for the companies doing it.

What the Next Decade Holds

The trajectory of AI drug discovery over the next decade points toward a world in which the boundaries between computational and experimental biology effectively dissolve. Not because laboratory experiments become unnecessary — they do not — but because the relationship between computational prediction and experimental validation becomes so tight and iterative that the two are essentially integrated aspects of a single research process.

The self-driving laboratory will become standard in high-throughput drug discovery programmes, not exceptional. The protein structure problem, solved by AlphaFold for static structures, will be extended to dynamic structures and protein-protein interactions, further deepening the precision of computational drug design. Multimodal AI systems that integrate genomic, proteomic, imaging, electronic health record, and wearable sensor data will identify disease subtypes and patient populations for clinical trials with unprecedented precision, improving the probability that the right drug reaches the right patient in the right trial.

The most ambitious vision — drug discovery that is not just AI-assisted but AI-driven, with AI systems setting research priorities, designing experiments, interpreting results, and proposing next steps with minimal human guidance — is probably 10 to 15 years away from being robust and reliable enough for routine application in complex disease areas. But in narrower, better-defined domains — structure-based drug design for validated targets with good assays — versions of this vision are already being built.

What remains irreducibly human in all of this? The judgment calls about which diseases to pursue. The interpretation of ambiguous clinical evidence. The ethical decisions about resource allocation. The communication with patients and clinicians about what drugs do and do not achieve. The relationship between science and the society it serves. These are not engineering problems. They are human responsibilities, and they do not get easier because the molecules are designed by machines.

Explore

More articles