Skip to content
AI.info

Research

Interpreting GFlowNets for Drug Discovery: What probes can and cannot show

Overview Research area: Interpretability of generative machine learning models for molecular design, specifically Generative Flow Networks (GFlowNets) applied to drug discovery. Technical level: Advan

arXiv
2511.19264
Published
2025-11-24
Authors
Amirtha Varshini A S, Duminda S. Ranasinghe, Hok Hei Tam

AI summary

Overview

Research area: Interpretability of generative machine learning models for molecular design, specifically Generative Flow Networks (GFlowNets) applied to drug discovery.

Technical level: Advanced. The paper assumes familiarity with sparse autoencoders, linear probing, graph transformers, and attribution methods, though the central negative control result is stated in plainly understandable terms.

Scope: The paper builds an interpretability toolkit for SynFlowNet, a GFlowNet trained on chemical reactions, and reports which interpretability claims survive rigorous controls — concluding that linear decodability of physicochemical properties reflects graph architecture and atom featurization rather than anything the trained policy learned.

What This Paper Is About

GFlowNets can generate molecules sequentially and allocate probability mass across diverse high-reward structures, and SynFlowNet constrains this generation to synthetically accessible chemical space. The problem is that these models' internal decision policies are opaque, and medicinal chemists need structural rationales before trusting proposed molecules. The authors set out to probe SynFlowNet's internal embeddings with saliency methods, sparse autoencoders, and motif probes, and then to test rigorously whether what those probes reveal is genuinely learned chemistry or merely information available for free in the molecular graph.

Key Contributions

  1. An integrated interpretability framework for GFlowNets spanning three levels of explanation: gradient-based saliency with motif-level counterfactual QED edits, sparse autoencoder concept attribution, and motif probes that test whether functional groups are decodable from pooled embeddings.

  2. A matched overcomplete-versus-undercomplete SAE comparison using an overcomplete (256 → 1024) BatchTopK sparse autoencoder against a matched linear baseline (256 → 128), showing the overcomplete model reconstructs the embedding 1.75 times better at comparable average sparsity on identical splits and seeds.

  3. A systematic control suite for linear probes. Eleven targets are probed on the raw embedding, the sparse latent, and the reconstruction, and validated against descriptor baselines, random-label controls, Bemis-Murcko scaffold-disjoint splits, and an architecture-matched untrained network. Single-feature ablation and cross-seed stability checks are added at the feature level.

  4. A decisive negative control. A randomly initialized network of identical architecture (GraphTransformerSynGFN, 4 layers, 2 heads, Morgan-1024 features, 2.9M parameters) decodes every probed property essentially as well as the trained policy, which the authors use to bound what probe scores can legitimately be claimed to show.

Main Findings

  • The sparse bottleneck preserves most probe-detectable information. Passing test embeddings through the encode-decode pipeline and re-applying pre-trained linear probes gives a median retention of 98.3%. Classification is essentially untouched (change ≤ 0.001 for halogen, aromaticity, and boron). The worst case is flexibility, which loses ΔR² = −0.116 from a base of 0.694, a 17% relative drop (83% retained).

  • One target improves through the bottleneck. Drug-likeness is decoded better from the sparse latent than from the raw embedding (ΔR² = +0.033), meaning the sparse basis makes QED more linearly accessible than the dense embedding does.

  • The overcomplete SAE is an exact-sparsity decomposition with negligible feature death. Reconstruction NMSE is 0.00151 ± 0.00009 with a batch-mean of exactly 64 active features per example (L₀ = 64.000 ± 0.000 across seeds; per-molecule counts vary) and feature death of 0.64% ± 0.28%. The matched linear baseline reaches NMSE 0.00263 ± 0.00011, 1.75 times worse, while reaching batch-mean L₀ = 67.6 on its own.

  • Absolute probe scores are high, but descriptors match or beat them. The embedding decodes QED at R² = 0.59, size at 0.996, and halogen at AUROC 1.000, with random-label controls at chance (AUROC 0.49 to 0.52). Under a random split the embedding never exceeds the leakage-free RDKit descriptor baseline by more than +0.036 (the complexity proxy); next-largest advantages are molecular weight and size (+0.016 each) and polarity (+0.008). On four targets the descriptor baseline is higher, by 0.026 (QED), 0.073 (lipophilicity), 0.077 (urea), and 0.202 (flexibility).

  • Scaffold splitting barely moves the embedding but does move the descriptors. Under a Bemis-Murcko scaffold-disjoint split (18,650 distinct scaffolds; 3,206 test molecules; zero train-test scaffold overlap; five splits), the largest embedding change is flexibility (−0.057) and drug-likeness moves by +0.001. The QED descriptor score falls from 0.612 to 0.568, flipping the gap from −0.026 to +0.019. Under the harder split the embedding loses on three of eleven targets (flexibility, urea, lipophilicity), is ahead on six, and is exactly tied on two (halogen and aromaticity, both 1.000 versus 1.000).

  • Training contributes almost nothing to linear decodability. The untrained network decodes every property to within 0.036 of the trained policy. On drug-likeness, training buys +0.005 (0.581 ± 0.016 untrained versus 0.586 trained), well inside the seed spread. The largest gain from training is flexibility at +0.036, and on molecular size and weight the untrained network is marginally better.

  • Chemically enriched substructure detectors emerge in individual runs. Across 3,959 enrichment tests, 280 are significant (gap above 30 points, q < 0.01). The strongest are single-halogen detectors: separate features fire almost only on iodine-, chlorine-, bromine-, and fluorine-containing molecules (100% versus 0 to 17% in silent molecules, q < 10⁻⁹). One feature detects boron and boronic esters (90% versus 0%, q = 5 × 10⁻¹¹), notable because boron appears in only 0.3% of the corpus.

  • The dictionary encodes structural absence as well as presence. One feature fires on molecules lacking a benzene ring (benzene present in 0% of its top activators versus 93% of silent molecules, q = 7 × 10⁻¹²).

  • Feature ablation effects are property-specific. Specificity ratios are roughly 8 to 23 times the off-target effect for the substantial interventions, with one small intervention reaching 44 times (targeted ΔR² ≈ −0.014 against an off-target mean near 0.0003). The most substantial interventions sit in the 8 to 23 times band with large targeted effects (for example ΔR² ≈ −1.67 at 23 times and −2.20 at 18 times). This is an interventional effect at the level of probe decoding; generation was never re-run.

  • Individual features are mostly not reproducible across seeds. Matching decoder columns across independently seeded runs gives a mean best-match cosine of 0.237 against a chance baseline of 0.214, only 1.11 times chance. A small core is real: 2.71% of features exceed cosine 0.5 and 0.97% exceed 0.7, thresholds no random pair attains.

  • Counterfactual saliency produces small but nonzero reward shifts. Integrated Gradients on the Stop log-probability highlights chemically meaningful regions, and RDKit-based edits within selected motifs shift QED by ΔQED = +0.015 and +0.051 in the two illustrative examples. The authors note these effect sizes are small relative to QED's range and that a controlled comparison against edits to randomly chosen motifs remains future work.

  • The analyzed checkpoint no longer exists. All analyses use frozen graph-level embeddings from a single SynFlowNet checkpoint trained with QED reward; that checkpoint is unavailable, so the extracted embeddings (32,054 molecules) are the surviving reproducible artifact and retraining would not reproduce the same weights. The two reported ΔQED values could not be independently re-verified.

Methodology in Plain English

The authors treat a trained molecular generator as a black box with an internal numerical summary (an embedding) for each molecule it considers, and they ask what information that summary contains and where that information comes from.

They gather frozen embeddings for 32,054 molecules from a SynFlowNet policy trained with QED as its reward function. Because the original checkpoint is gone, these embeddings plus the code are the artifact everyone can work from.

Three kinds of interrogation follow. First, gradient-based saliency with integrated gradients estimates which atoms drove the decision to stop generating at the final step, using M = 64 interpolation steps. High-scoring atoms are grouped into connected components and ring systems (RDKit's GetSymmSSSR) to form candidate motifs, and six fixed chemically motivated RDKit transformations (ether to thioether, methyl to fluorine, chloro to bromo, amide to ester, and others) are applied to those motifs to see how computed QED changes.

Second, two sparse autoencoders decompose the embedding. An undercomplete four-layer MLP autoencoder compresses 256 dimensions to a 128-dimensional sparse code with an ℓ1 penalty, optimized with Adam for 200 epochs at batch size 128 under a random_state of 42, and its factors are correlated against descriptors such as TPSA and Crippen logP. An overcomplete BatchTopK autoencoder expands 256 dimensions to 1024 with k = 64, trained for 50 epochs at batch size 512 with a 90/10 split and five seeds (42 to 46), with dead features resampled every 10 epochs from epoch 20 and capped at 2% per round.

Third, motif probes and linear probes test decodability. Motif probes are shallow three-layer MLPs (256 → 256 → 128 → 64 → 1) trained with class-balanced sampling for 50 epochs to predict SMARTS-defined functional groups from the embedding. Linear probes are Ridge and logistic regressions run on the raw embedding, the sparse latent, and the reconstruction across eleven targets, evaluated with an 80/10/10 split and probe seeds 0 to 4.

The distinguishing feature of the work is the control design. Each probe is run against shuffled labels, against an RDKit descriptor plus ECFP4 baseline with no model involved and with leakage-prone columns dropped per target, against a scaffold-disjoint split, and against an untrained network of identical architecture. Individual features are ablated one at a time to test specificity, and dictionaries from different seeds are matched by cosine similarity against a random-dictionary baseline.

Why This Matters

The paper's central message is a methodological caution: high linear-probe scores on a molecular generative model are not evidence that the model learned chemistry, because an untrained network with the same architecture and atom featurization scores just as well. This reframes how interpretability results for these models should be reported and which claims should be made.

Impact on research:

  • It supplies a concrete control protocol (random-label, descriptor, scaffold-split, untrained-network, ablation, cross-seed stability) that other interpretability studies of molecular generative models can adopt.
  • It demonstrates that overcomplete sparse dictionaries are a better reconstruction tool than a matched compressive baseline at comparable sparsity, while also showing that individual feature indices are not portable across runs.
  • It shows that probe decodability of physicochemical properties should be attributed to graph architecture and input featurization, not to policy training, which changes what benchmarks in this area can claim.
  • It documents the practical cost of non-reproducible checkpoints: because the analyzed weights no longer exist, certain results (the reported ΔQED values) cannot be re-verified.

Real-world applications (as directions the work speaks to, not as demonstrated deployments):

  • Medicinal chemists evaluating a generative model's proposed structures and synthetic routes, who need to know which rationales are trustworthy.
  • DMTA (design–make–test–analyze) pipelines that would need model explanations before integrating generative proposals.
  • Model auditing and validation in industrial drug discovery, where interpretability claims influence adoption decisions.
  • Feature-level analysis of learned representations using sparse autoencoders, which the paper shows yields per-halogen and boron detectors and an explicit absence detector.

Industry relevance: All three authors are affiliated with Montai Therapeutics, and the framing is explicitly industrial — it addresses the trust gap that keeps generative models out of chemists' workflows and states plainly which claims the evidence does not support.

Future Directions

  1. An architecture-matched untrained control under the scaffold-disjoint split. The authors ran the untrained control only under the random split and state that the same conclusion should hold under the scaffold split, but that this is untested.

  2. A controlled saliency-versus-random comparison. The counterfactual edits to high-attribution motifs produce small QED changes, and the authors state that comparing against edits to randomly chosen motifs is needed to establish the relationship quantitatively.

  3. Saliency at intermediate decision points. The current saliency analysis only covers the Stop action, the final decision step where the whole molecule is present; extending attribution to earlier action probabilities in the trajectory is named as important future work.

  4. Better targets for positive claims. The one target where the embedding leads the descriptor baseline is a clipped, narrow-range complexity proxy; the authors flag a genuine synthetic-accessibility or Bertz-complexity target as a better test. They also note that steering and generation-time controllability claims are entirely open, since generation was never re-run.

Target Audience

Machine learning researchers working on interpretability of generative models, particularly those studying sparse autoencoders, linear probing, and attribution methods on non-language domains. Computational chemists and cheminformatics practitioners who use or evaluate generative molecular design tools and need to judge whether interpretability claims are well-founded. Industry scientists in drug discovery responsible for model validation, auditing, and deciding whether a generative model is ready to enter design–make–test–analyze workflows. Readers specifically interested in negative results and in what controls are necessary before attributing learned knowledge to a trained model will find the paper most valuable.

Authors’ abstract

Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures. We present a control-validated interpretability study of SynFlowNet, a synthesis-aware GFlowNet trained with a drug-likeness (QED) reward. Our framework combines gradient saliency and counterfactual edits, an undercomplete factor analysis, and an overcomplete BatchTopK sparse autoencoder, evaluated with shuffled-label controls, RDKit-descriptor baselines, a scaffold-disjoint split, cross-seed stability, and an architecture-matched untrained network. These controls materially change the interpretation. Physicochemical properties and functional groups are highly decodable from SynFlowNet embeddings, but an untrained network with the same architecture performs essentially as well as the trained policy (drug-likeness within 0.005 and marginally better molecular-size decoding). Thus, this decodability reflects graph architecture and atom featurization rather than representations acquired through policy training. The surviving conclusions are narrower: the overcomplete autoencoder reconstructs embeddings better than a matched undercomplete baseline at comparable sparsity; chemically enriched substructure detectors, including per-halogen and boron features, emerge in individual runs, although only a small subset of dictionary directions is stable across seeds; and zeroing individual features produces property-specific effects on probe decoding. Beyond SynFlowNet, this study provides a reusable control protocol for molecular-model interpretability, separating learned structure from signals supplied by architecture and input representation. High probe scores alone should not be treated as evidence of learned chemistry.

Read the original paper