Skip to content
AI.info

Research

From In Silico to In Vitro: Evaluating Molecule Generative Models for Hit Generation

Overview Research area: Machine learning for drug discovery — specifically deep generative models applied to the hit identification stage of the pharmaceutical pipeline. Technical level: Intermediate.

From In Silico to In Vitro: Evaluating Molecule Generative Models for Hit Generation
arXiv
2512.22031
Published
2025-12-26
Authors
Nagham Osman, Vittorio Lembo, Giovanni Bottegoni, Laura Toni

AI summary

Overview

Research area: Machine learning for drug discovery — specifically deep generative models applied to the hit identification stage of the pharmaceutical pipeline.

Technical level: Intermediate. The paper assumes familiarity with molecular graph generation, autoregressive vs. diffusion models, and structure-based docking, but explains its task framing and metrics clearly.

Scope: A task-focused benchmark of three graph-based molecule generators (MolRNN, GraphINVENT, DiGress) against a newly defined "hit-like" chemical space, combining medicinal-chemistry filters, docking-score distributions across seven protein targets, and prospective in vitro validation of three synthesized GSK-3β candidates.

What This Paper Is About

Most molecular generative models are judged on whether they produce molecules that are valid and drug-like in general, not on whether they can do the specific job of finding a hit — the first small molecule with reproducible activity, acceptable synthetic accessibility, and suitable physicochemical properties that enters the drug discovery pipeline. This paper reframes the question as: can off-the-shelf generative models be trained to directly generate hit-like molecules, substituting for or feeding into traditional hit identification workflows such as high-throughput screening and virtual screening? The authors build an explicit hit-likeness definition, benchmark three models under multiple training regimes, and then go as far as synthesizing and testing the best candidates in the lab.

Key Contributions

  1. A task definition and evaluation framework for hit generation. The authors operationalize "hit-likeness" as a standalone downstream task, combining physicochemical, structural, and bioactivity-related criteria — including distributional alignment to known ligands and target-specific docking distributions — rather than proposing a new scoring function or optimization objective. They state this is the first task-oriented, end-to-end evaluation of off-the-shelf generative models explicitly focused on hit generation.

  2. A multi-stage hit-like filtering pipeline. The framework integrates established drug-likeness and medicinal-chemistry constraints (Novartis-inspired severity score, molecular weight, logP, pChEMBL activity, synthetic accessibility, ring-system rules, and element restrictions) to define and evaluate the hit-like chemical space, and to measure which criteria each model most often fails.

  3. A benchmark of two autoregressive models and one diffusion model across diverse datasets and training settings. MolRNN, GraphINVENT, and DiGress were trained on the full REINVENT dataset, a newly constructed hit-like dataset, and a fine-tuning regime combining the two, with MolRNN and DiGress additionally fine-tuned on seven target-specific ligand sets. Outputs were assessed with standard metrics (validity, uniqueness, novelty, FCD, fragment/scaffold similarity, nearest-neighbor similarity, internal diversity) plus target-specific docking scores.

  4. Prospective in vitro validation. Three selected GSK-3β candidates were synthesized and tested, with one compound showing low-nanomolar inhibition, demonstrating that generative models can produce biologically meaningful hits, not just chemically valid structures.

Main Findings

  • Validity and novelty differ sharply by architecture. MolRNN and GraphINVENT achieved near-perfect validity (>99.9%) across all settings, while DiGress reached only 70–76%. Uniqueness was uniformly high. Novelty was highest for GraphINVENT (consistently >93%), sensitive to training data for MolRNN (dropping to 64.60% in the Hit-like regime), and moderate for DiGress (~70–75%).

  • Hit-like filter compliance depends heavily on training data. Direct training on the hit-like dataset yielded the highest compliance (MolRNN: 76.87%), while REINVENT-trained models passed at under 12%. MolRNN and GraphINVENT improved substantially with fine-tuning (MolRNN 67.57%, GraphINVENT 50.80%), but DiGress saw only modest gains (14.11%) and remained poor even when trained directly on hit-like data.

  • The most common filter failures are size and lipophilicity. MolRNN (REINVENT) and DiGress (REINVENT) frequently exceeded molecular weight and logP thresholds. GraphINVENT tended toward higher ring-complexity and logP failures, partly mitigated by fine-tuning. DiGress overrepresented large and fused ring systems, with high MW violations persisting after fine-tuning.

  • Diffusion-based generation struggled in the constrained regime. DiGress consistently underperformed autoregressive models in validity and filter compliance. The authors attribute this to discrete diffusion relying on global denoising, making it more sensitive to distributional sparsity and strict post hoc filters, whereas autoregressive models enforce chemical validity at each generation step.

  • Distributional similarity favored fine-tuned autoregressive models. GraphINVENT (Hit fine-tune) achieved the lowest FCD (1.13), while MolRNN showed the highest scaffold similarity, and DiGress generally showed higher FCD and scaffold deviation.

  • Docking-score distributions aligned best with hit-like training. Models trained directly on the hit-like set achieved the lowest KL divergence against reference ligand sets (all <0.01): GraphINVENT (Hit-like) 0.003 ± 0.002, MolRNN (Hit-like) 0.004 ± 0.003, DiGress (Hit-like) 0.009 ± 0.012. REINVENT-trained models ranged 0.083–0.100, and fine-tuned models were intermediate. Averaged by algorithm, MolRNN was lowest (0.061 ± 0.062), with GraphINVENT (0.064 ± 0.067) and DiGress (0.066 ± 0.062) similar.

  • Target difficulty varies. Per-target analysis showed the highest KL divergences for SRC and PPARα, reflecting the higher molecular weight and logP of their known ligands relative to the generated set. GSK-3β and ADORA2A showed the closest alignment with hit-like constraints. Across the 63 model–training combinations, HSP90α showed the highest docking success rates (38–52%), PPARα the lowest (<1%), and GSK-3β moderate (10–16%).

  • Restricting the reference set changes the picture. Repeating the docking comparison using only ligands that pass the hit-like filters raised success rates for five targets, most notably thrombin (from <2% to 35–45%), while GSK-3β and HSP90α observed modest drops.

  • MolRNN fine-tuning was the strongest performer overall. MolRNN (Hit fine-tune) achieved the highest docking success in 6 of 7 targets, with DiGress (Hit fine-tune) narrowly leading for PPARα. The weakest performances came from DiGress (Hit-like) in four targets, GraphINVENT (REINVENT) in two, and GraphINVENT (Hit-like) for HSP90α.

  • Experimental validation succeeded for one compound. Compound 1 inhibited GSK-3β with an IC50 of 314 ± 15.3 nM, showing strong inhibition already at 1 µM (71.3 ± 1.7%; 88.2 ± 0.1% at 10 µM; 89.9 ± 0.2% at 50 µM). Compound 2 showed moderate inhibition at 50 µM (49.6 ± 0.9%) and compound 3 weak inhibition (5 ± 0.7%); compounds 2 and 3 were not inhibited at the lower concentrations.

  • The validated hit is chemically novel. Compound 1's 1H-pyrazolo[3,4-b]pyridine scaffold is rare among GSK-3β inhibitors (19 of 1734, 1.1%) and among all kinase inhibitors in ChEMBL (727 of 100365, 0.72%). Docking analysis showed hinge-region binding through two hydrogen bonds (Asp133, Val135) plus a Lys85 interaction. Its –log(activity) of 6.50 exceeded the median of the full ligand set (6.39) and the hit-like inhibitor subset (5.89).

  • Standard metrics did not track bioactivity. The authors report that widely used metrics such as VUN, FCD, and scaffold similarity did not consistently correlate with predicted bioactivity.

Methodology in Plain English

The authors took three graph-based generative models that perform well on general molecule generation — MolRNN (a GRU-based autoregressive model), GraphINVENT (an autoregressive model using a message-passing neural network), and DiGress (a discrete diffusion model with a graph transformer) — and asked whether they can be repurposed for hit generation.

They assembled three kinds of data. A general-purpose set (the REINVENT dataset, 1,086,248 unique compounds filtered for 10–50 heavy atoms and a fixed element list), a hit-like set they constructed by applying their hit-like criteria to the same ChEMBL source (58,837 molecules, roughly 2.68% of the original database), and target-specific ligand sets for seven proteins (D3R, ADORA2A, HSP90α, GSK-3β, SRC, thrombin, and PPARα), selected for pChEMBL value ≥ 5 and assay confidence score = 9.

Each model was trained under three settings: on the full REINVENT dataset, on the hit-like dataset, and by fine-tuning a REINVENT-trained model on the hit-like dataset. MolRNN and DiGress were additionally fine-tuned on the seven target-specific sets. All models followed their original training objectives with early stopping on validation loss; fine-tuning ran up to 20 epochs with smaller batch sizes. The best checkpoint per model was selected by lowest validation loss, and generation continued until each model produced the same number of VUN- and hit-like-filtered molecules, so that docking and drug-likeness comparisons were made on equal footing.

Evaluation proceeded in three layers of increasing drug relevance. The first layer covered standard validity, uniqueness, and novelty checks. The second measured resemblance to the training chemical space using Fréchet ChemNet Distance, BRICS fragment and Bemis–Murcko scaffold similarities, nearest-neighbor Tanimoto similarity, and Jaccard-based internal diversity on Morgan fingerprints. The third layer applied the hit-like filters, then computed docking scores only for molecules passing both VUN and the full filter set, using Glide (Schrödinger 2023–04, SP protocol, OPLS4 force field) across the seven targets.

For biological evaluation, GSK-3β was chosen because models showed low KL divergence there and many compounds beat the median docking score of known ligands. Candidates had to be structurally novel (Tanimoto distance ≥ 0.5 from known binders) and strongly predicted binders (docking scores below mean − 2 SD), then were prioritized by docking score, visual inspection of binding poses, and synthetic feasibility. Three compounds were synthesized and tested with a LANCE Ultra TR-FRET GSK-3β kinase assay, screened at 1, 10, and 50 µM, with the most potent compound taken into an 11-concentration dose–response from 300 pM to 100 µM.

Why This Matters

The paper moves the conversation from "can generative models make valid molecules" to "can they do a specific job in the drug discovery pipeline." That shift matters because validity and drug-likeness metrics have repeatedly failed to predict whether a molecule will actually bind a target — a gap this study documents directly by showing that VUN, FCD, and scaffold similarity did not track predicted bioactivity, and by having a generated, structurally novel compound confirmed active at nanomolar potency in vitro.

Real-world applications:

  • Hit identification acceleration. Teams can replace or augment high-throughput screening campaigns that take months to years and are financially demanding with generative models trained on hit-like data, then triage outputs through docking before committing to synthesis.

  • Target-specific candidate generation. The target-specific fine-tuning experiments (seven targets spanning GPCRs, kinases, a protease, a chaperone, and a nuclear receptor) show a route to generating compounds for a defined biological target, useful for programs where no large ligand set exists.

  • Benchmark and dataset design. The finding that off-the-shelf models trained on broad chemical libraries fail hit-like filters at rates above 87% is a concrete argument for curating hit-like training sets and for evaluation protocols that reward biological relevance over distributional similarity.

  • Lead optimization handoff. Because the best molecules already satisfy stringent physicochemical and synthetic-accessibility criteria, they can feed directly into hit-to-lead optimization, whether run by ML models or medicinal chemists.

Industry relevance: The paper targets the earliest, most expensive, and most failure-prone stage of pharmaceutical discovery. The prospect of compressing a multi-year process into a few months is explicitly framed in the introduction, and the results give a realistic picture of which model classes and training regimes actually deliver compounds worth synthesizing.

Future Directions

  • Build richer, better-curated hit-like datasets. The authors attribute constrained performance on PPARα and SRC to scarce ligands with suitable physicochemical properties, calling for better data even for well-studied proteins, to support both general and target-specific training.

  • Fix the metrics. Because VUN, FCD, and scaffold similarity did not consistently correlate with predicted bioactivity, evaluation frameworks need to integrate bioactivity-relevant endpoints alongside distributional and structural measures.

  • Improve architectures for low-data, target-specific regimes. Diffusion models struggled with target-specific fine-tuning due to higher data demands and sensitivity to dataset size, pointing to improved transfer learning or data augmentation as key directions.

  • Assemble a modular AI-driven pipeline. The authors propose separating concerns: generative models focus on high-quality hit generation, followed by ML-based or traditional hit-to-lead optimization, rather than asking one model to cover the whole pipeline.

Target Audience

This paper is most useful to machine learning researchers working on molecular generation who want a realistic benchmark beyond MOSES and GuacaMol-style validity metrics; to computational chemists and medicinal chemists who need to know whether generated molecules are worth synthesizing; and to drug discovery teams evaluating generative models for early-stage hit identification. It also serves as a methodological reference for researchers designing task-specific, experiment-backed evaluations of generative models in scientific domains beyond chemistry.

Authors’ abstract

Hit identification is a critical yet resource-intensive step in the drug discovery pipeline, traditionally relying on high-throughput screening of large compound libraries. Despite advancements in virtual screening, these methods remain time-consuming and costly. Recent progress in deep learning has enabled the development of generative models capable of learning complex molecular representations and generating novel compounds de novo. However, using ML to replace the entire drug-discovery pipeline is highly challenging. In this work, we rather investigate whether generative models can replace one step of the pipeline: hit-like molecule generation. To the best of our knowledge, this is the first study to explicitly frame hit-like molecule generation as a standalone task and empirically test whether generative models can directly support this stage of the drug discovery pipeline. Specifically, we investigate if such models can be trained to generate hit-like molecules, enabling direct incorporation into, or even substitution of, traditional hit identification workflows. We propose an evaluation framework tailored to this task, integrating physicochemical, structural, and bioactivity-related criteria within a multi-stage filtering pipeline that defines the hit-like chemical space. Two autoregressive and one diffusion-based generative models were benchmarked across various datasets and training settings, with outputs assessed using standard metrics and target-specific docking scores. Our results show that these models can generate valid, diverse, and biologically relevant compounds across multiple targets, with a few selected GSK-3$β$ hits synthesized and confirmed active in vitro. We also identify key limitations in current evaluation metrics and available training data.

Read the original paper