Research
MolGuidance: Advanced Guidance Strategies for Conditional Molecular Generation with Flow Matching
Overview Research area: Machine learning for scientific discovery — specifically conditional 3D molecular generation using SE(3)-equivariant flow matching, evaluated on quantum-chemistry datasets. Tec

- arXiv
- 2512.12198
- Published
- 2025-12-13
- Authors
- Jirui Jin, Cheng Zeng, Pawan Prakash, Ellad B. Tadmor, Adrian Roitberg, Richard G. Hennig, Stefano Martiniani, Mingjie Liu
AI summary
Overview
- Research area: Machine learning for scientific discovery — specifically conditional 3D molecular generation using SE(3)-equivariant flow matching, evaluated on quantum-chemistry datasets.
- Technical level: Advanced. The paper assumes familiarity with flow matching, continuous-time Markov chains (CTMCs), classifier-free guidance, and equivariant graph neural networks, though its central conclusions are stated in plain comparative terms.
- Scope (1 sentence): The paper integrates and systematically benchmarks several guidance strategies — classifier-free guidance (CFG), autoguidance (AG), and model guidance (MG) — into a flow-matching molecule generator, proposes a hybrid scheme that guides continuous and discrete molecular features separately, and reports state-of-the-art property alignment on QM9 and QMe14S.
What This Paper Is About
Generating molecules with precisely controlled properties is hard because a molecule is not a single data type: it combines discrete variables (atom types, formal charges, bond orders) with continuous ones (3D atomic coordinates), and guidance methods borrowed from image generation do not automatically transfer to this multi-modal setting. The authors build on their existing property-conditioned flow-matching model, PropMolFlow (PMF), to test how different guidance strategies and mathematical formats should be applied to each molecular modality. Their goal is a rigorous comparison of guidance methods and a hybrid strategy that achieves better property alignment than prior conditional generation models.
Key Contributions
- A unified guidance framework for multi-modal molecules (MolGuidance). The authors integrate three advanced guidance strategies — CFG, autoguidance, and model guidance — with an SE(3)-equivariant flow-matching molecule generator, and additionally apply guidance separately to continuous modalities (velocity fields for atomic positions) and discrete modalities (predicted logits/probabilities for atom types, charges, and bonds).
- The first systematic study of discrete guidance formats. The paper compares four formulations — Linear-Prob, Log-Prob, Linear-Rate, and Log-Rate — and reports that the empirically motivated linear-on-probability and logarithmic-on-probability forms outperform the theoretically derived logarithmic-on-rate-matrix form, which becomes numerically unstable.
- A hybrid guidance strategy optimized with Bayesian optimization. Separate guidance weights for continuous (w₁) and discrete (w₂) features are jointly tuned by Bayesian optimization, which the authors report as pushing state-of-the-art property alignment for de novo molecular generation.
- Benchmarks and DFT validation across two datasets. The work reports results on QM9 (5 elements) and the larger, more element-diverse QMe14S (14 elements), including density functional theory checks on generated molecules.
Main Findings
- Logarithmic guidance on rate matrices is unstable. While theoretically derived by Nisonoff et al., this format shows severe instability at discrete guidance weights beyond w₂ = 1.2 and was constrained to the narrow range 1.0–1.2 in increments of 0.05, whereas the other three formats were explored from 1.0 to 3.0 in increments of 0.5.
- Logarithmic guidance on probability distributions gives the best alignment with stability. It reaches its minimum property MAE for dipole moment around w₂ = 2.5 while remaining numerically stable across the full weight range; linear formats improve more gradually and need higher guidance strengths for comparable performance.
- Discrete guidance beats continuous guidance at sampling time. Discrete-only guidance (atom types, charges, bonds) substantially outperforms continuous-only guidance (atomic positions), which contradicts the training loss hierarchy in which atomic positions receive the dominant weight — revealing that training and guided sampling pose different optimization challenges.
- Hybrid guidance is best. Combining continuous and discrete guidance consistently achieves superior property alignment on both datasets for dipole moment, with optimal weights typically between w = 2.0 and 3.0 for numerically stable methods.
- Bayesian-optimized CFG reaches the best alignment. For CFG conditioned on dipole moment, optimal performance (MAE = 202 meV) occurs at (w₁, w₂) = (2.71, 1.91), and alignment depends more on the discrete weight w₂ than on the continuous weight w₁.
- PMF-CFG achieves the best alignment for four of six properties. It leads on polarizability (α), HOMO-LUMO gap (Δε), ε_HOMO, and dipole moment (μ), while remaining competitive on ε_LUMO and heat capacity (C_v) against JODO. Example: MAE of 1.27 Bohr³ on α versus 1.97 Bohr³ for the GCDM model.
- PMF-AG is the most balanced method. It surpasses PMF-Vanilla across all properties, outperforms JODO on α, matches it on Δε, ε_HOMO, ε_LUMO, and μ, and falls slightly behind on C_v.
- Model guidance underperforms in this domain. PMF-MG models underperform their vanilla counterparts across all properties, which the authors attribute to the difficulty of jointly learning property constraints and guidance-scale embeddings.
- QMe14S MAEs are generally lower than QM9 MAEs. The relative ranking of guidance methods stays consistent across datasets despite the jump from 5 to 14 elements.
- DFT validation supports the alignment claims. Single-point B3LYP/6-31G(2df,p) calculations for QM9 and B3LYP/TZVP for QMe14S were run on 500 molecules selected from 10,000 generated molecules; DFT properties of relaxed molecules were used for C_v due to vibrational frequency issues. The authors note an underestimation of DFT MAEs against input target values for Δε, ε_HOMO, and ε_LUMO when a GVP property predictor is used.
- Structural validity remains high. All guidance methods outperform GeoLDM and GCDM on molecule stability, and edge out JODO, which uses explicit bond information. PMF-CFG incurs a 2–3.4% decline in molecule stability relative to the vanilla model for Δε, ε_LUMO, μ, and C_v, while α and ε_HOMO remain essentially unchanged. PMF-AG improves stability across all properties relative to PMF-Vanilla; MG models achieve the highest stability for Δε and μ.
- Diversity trades off against alignment. CFG shows the most notable decline in valid-and-unique rates relative to vanilla, followed by AG and then MG. CFG excels in element entropy and performs well across bond-order entropy and scaffold diversity; AG shows high bond-order and element entropy but falls short in scaffold diversity.
- Sampling cost differs by method. MG and PMF-Vanilla require a single forward pass and are fastest; CFG and AG require two forward passes, with AG faster than CFG because its guide network is much smaller than CFG's unconditional model.
- Two guidance weights are sufficient. Bayesian optimization with four separate weights (one per modality) yields performance comparable to using two weights (one for positions, one shared across discrete variables); similar results were observed for AG.
- More integration steps have limited benefit. Increasing inference time steps slightly reduces property MAEs for Δε, ε_HOMO, and ε_LUMO, but has negligible effect for the other three properties.
Methodology in Plain English
The authors start from PropMolFlow, their property-conditioned flow-matching model, which learns to transform Gaussian noise into molecular graphs conditioned on target properties. A molecule is represented as a fully connected graph G = (X, A, C, E) — atomic positions, atom types, charges, and bond orders — and each modality is evolved by its own process: continuous flow matching for positions, and a continuous-time Markov chain (CTMC) process for the discrete variables. The model backbone is an SE(3)-equivariant graph neural network (chosen over E(3) to preserve chirality by not forcing reflection symmetry).
They then add guidance in two independent places. For continuous positions, they linearly interpolate between unconditional and conditional velocity fields with weight w₁. For discrete variables, they test four options: linear interpolation of denoising probabilities, logarithmic interpolation of probabilities, linear interpolation of rate matrices, and logarithmic interpolation of rate matrices, each governed by weight w₂. They compare these formats by sweeping guidance weights and measuring the mean absolute error (MAE) between target property values and values predicted by a pretrained regressor on the generated molecules.
Next they run controlled experiments isolating continuous-only, discrete-only, and hybrid guidance to see which modality matters more during sampling. They then use Bayesian optimization to jointly tune the guidance weights for each method, and finally benchmark CFG, AG, and MG against a vanilla conditional model and three diffusion baselines (GeoLDM, GCDM, JODO) on property alignment, molecule stability, RDKit validity, uniqueness, and wall-clock sampling time. DFT calculations on 500 generated molecules provide an independent check on property values.
Why This Matters
This work gives practitioners a concrete, empirically tested playbook for applying guidance in multi-modal generative models for chemistry. It shows that mathematical elegance (the theoretically derived rate-matrix formulation) can lose to a simpler empirical formulation when numerical stability matters, and it overturns an intuitive assumption — that the modality receiving the most training emphasis should receive the most guidance at sampling time. The transferability from 5-element QM9 to 14-element QMe14S suggests the framework generalizes rather than overfitting to one chemical space.
Real-world applications:
- Drug design: generating candidate molecules that hit target electronic or thermodynamic property values before committing to synthesis.
- Materials discovery: applying the same multi-modal guidance machinery to inorganic and organic materials generation, where property targets are similarly scalar and chemistry is multi-modal.
- Property-targeted screening: using CFG's precise property alignment to build focused virtual libraries, or AG's balance when diversity and validity matter as much as alignment.
- Benchmarking pipelines: providing reference numbers and trade-off analyses for teams comparing guidance methods in their own generative chemistry models.
Industry relevance: computational chemistry and AI-driven drug discovery groups gain a directly usable recipe for conditioning generative models on quantum-mechanical properties, along with quantified trade-offs between alignment, validity, diversity, and sampling cost — the exact considerations that determine whether a generative model is deployable in a discovery pipeline.
Future Directions
- Extending to high-dimensional and multi-property conditions. The paper states that the current study focuses on a single scalar property; tensorial properties and joint multi-property conditioning require methodological development.
- Better guide models for autoguidance. The authors note they studied two strategies for constructing the guide model and suggest that further exploration, including EMA strategies, may improve performance.
- Beyond linear interpolation and Gaussian priors. Guidance based on linear interpolation of velocity fields may break on general Riemannian manifolds or with a non-Gaussian base distribution.
- Inference-time optimization. The authors point to recent work on adding control support to steer generation toward desired directions during sampling as a promising complementary direction.
- Transfer to other scientific domains. The guidance methods and formats are described as readily transferable to de novo generation for inorganic and organic materials as well as protein design.
Target Audience
This paper is most valuable to machine learning researchers and computational chemists working on generative models for molecular design — particularly those already familiar with diffusion or flow matching who want to add property conditioning to an existing model. It also serves practitioners in AI-driven drug discovery and materials informatics who need to choose a guidance strategy and understand its trade-offs, and method developers interested in how guidance techniques from computer vision behave when the data is multi-modal, equivariant, and chemically constrained. Readers without a background in generative modeling will find the technical sections dense, but the comparative results and trade-off analysis are accessible on their own.
Authors’ abstract
Key objectives in conditional molecular generation include ensuring chemical validity, aligning generated molecules with target properties, promoting structural diversity, and enabling efficient sampling for discovery. Recent advances in computer vision introduced a range of new guidance strategies for generative models, many of which can be adapted to support these goals. In this work, we integrate state-of-the-art guidance methods -- including classifier-free guidance, autoguidance, and model guidance -- in a leading molecule generation framework built on an SE(3)-equivariant flow matching process. We propose a hybrid guidance strategy that separately guides continuous and discrete molecular modalities -- operating on velocity fields and predicted logits, respectively -- while jointly optimizing their guidance scales via Bayesian optimization. Our implementation, benchmarked on the QM9 and QMe14S datasets, achieves new state-of-the-art performance in property alignment for de novo molecular generation. The generated molecules also exhibit high structural validity. Furthermore, we systematically compare the strengths and limitations of various guidance methods, offering insights into their broader applicability.