Research
MolSculpt: Sculpting 3D Molecular Geometries from Chemical Syntax
Overview Research area: Machine learning for molecular science — specifically generative modeling of 3D molecular geometries, combining a frozen 1D molecular foundation model with a 3D diffusion model

- arXiv
- 2512.10991
- Published
- 2025-12-09
- Authors
- Zhanpeng Chen, Weihao Gao, Shunyu Wang, Yanan Zhu, Hong Meng, Yuexian Zou
AI summary
Overview
- Research area: Machine learning for molecular science — specifically generative modeling of 3D molecular geometries, combining a frozen 1D molecular foundation model with a 3D diffusion model.
- Technical level: Advanced. The paper assumes familiarity with diffusion models, transformer conditioning mechanisms (adaLN), geometric graph representations, and molecular string representations such as SELFIES.
- Scope in one sentence: The paper introduces MolSculpt, a framework that extracts latent chemical knowledge from a frozen pretrained 1D molecular language model via learnable queries and injects it as conditioning into a 3D diffusion model, and evaluates it on de novo and property-conditional 3D molecule generation on GEOM-DRUGS and QM9.
What This Paper Is About
Generating valid 3D molecular structures is hard because 1D sequence models (which guarantee chemical validity through SELFIES) and 3D geometry generators (which produce coordinates) have traditionally operated in isolation, so the rich chemical knowledge learned by 1D models is never used to guide 3D generation. The authors build a bridge between the two: a set of learnable queries probes a frozen 960M-parameter molecular foundation model (MoLLaMA) and a trainable projector maps the extracted information into the conditioning space of a diffusion model. The goal is to produce 3D conformers that are simultaneously syntactically valid and geometrically accurate.
Key Contributions
- Cross-modal knowledge transfer for 3D molecule generation. MolSculpt uses learnable query tokens to extract inherent chemical knowledge from a frozen 1D molecular foundation model (MoLLaMA) and conditions a 3D diffusion model on it, rather than using only 3D features or heuristic multimodal fusion as prior methods do.
- A trainable projector with bi-directional attention. Because the learnable queries have no inherent sequential ordering, the projector (built on the Qwen2.5 transformer encoder architecture) allows each query token to attend to all others and to the SELFIES context, producing compact chemical priors that are mapped through an FFN into the diffusion model's conditioning interface.
- A two-stage training recipe that keeps the foundation model frozen. Stage 1 trains 1D molecule generation (100 epochs on QM9-2014, 20 epochs on GEOM-DRUGS following Liu et al. 2025); Stage 2 trains the projector, FFN, and diffusion model end-to-end against the original diffusion objective while MoLLaMA remains frozen, preserving generalizable priors and avoiding distribution sharpening.
- State-of-the-art claims across two benchmarks. The paper reports leading results in de novo 3D molecule generation on both GEOM-Drugs and QM9, and achieves the lowest MAE on all six evaluated quantum-chemical properties for conditional generation on QM9.
Main Findings
- De novo generation on GEOM-Drugs (3D metrics): MolSculpt achieves an FCD_3D of 13.67, better than all baselines and below the training-set upper bound of 13.73 (NExT-Mol 14.69, UDM-3D 17.36). It attains a 3D MolStable of 0.026 (training set 0.028) and 3D AtomStable of 0.854 (training set 0.861).
- De novo generation on GEOM-Drugs (2D metrics): FCD of 0.330, MolStable 0.999, AtomStable 1.000, V&C 1.000, V&U 0.999, V&U&N 0.947, SNN 0.531, Frag 0.999, and a Scaf score of 0.605 — the highest reported, above the training set's 0.584.
- De novo generation on QM9 (2D metrics): SOTA results reported for FCD (0.065), MolStable (0.999), V&C (1.000), SNN (0.533), and Scaf (0.945), with AtomStable matching the best at 1.000.
- De novo generation on QM9 (3D metrics): MolSculpt reports a SOTA Bond Angle MMD of 2.90e-03, the best 3D AtomStable (0.995) and 3D MolStable (0.961). Its FCD_3D of 0.952 is noted as slightly behind the best baseline, NExT-Mol at 0.879.
- Conditional generation on QM9 (MAE): MolSculpt achieves the lowest MAE on all six properties — μ(D) 0.338, α(Bohr³) 0.29, C_v (cal/mol·K) 0.131, ε_HOMO 63 meV, ε_LUMO 62 meV, and Δε 85 meV. Relative to the runner-up NExT-Mol, this corresponds to reductions of 33.3%, 75.0%, 74.4%, 69.3%, 73.6%, and 71.4% respectively.
- Architecture ablation: Removing MoLLaMA ("Stage 2 w/o MoLLaMA") worsens FCD_3D from 0.952 to 0.998 and lowers molecular stability; using a fine-tuned MoLLaMA in Stage 2 also yields a worse FCD_3D of 0.969, supporting the decision to keep it frozen.
- Query token ablation: FCD_3D is 0.972 with 16 tokens, 0.976 with 32, an optimal 0.952 with 64, then 0.962 with 128 and 0.964 with 256 — the authors attribute the drop at higher counts to redundant information or noise, and set N_Q = 64 as default.
- Projector depth ablation: A 0-layer projector gives FCD_3D of 1.233, 6 layers 0.977, 12 layers 0.954, and 24 layers 0.952. The authors adopt 24 layers while noting the gain over 12 layers is marginal (performance saturates).
- Geometric distribution analysis: Generated bond length, bond angle, and dihedral angle distributions align closely with the test set, including multi-modal dihedral distributions for C-C-C-C in QM9 and c-c-c-H in GEOM-Drugs. The main discrepancy is that C-H bond lengths are slightly broader and less peaked than the real distribution, which the authors describe as a minor excess of flexibility in otherwise rigid bonds.
- Qualitative results: The authors report that generated molecules are chemically valid, structurally diverse, and that on the complex GEOM-Drugs dataset MolSculpt avoids artifacts such as disconnected components and preserves the planarity of aromatic rings.
Methodology in Plain English
The approach starts from a geometric graph representation of a molecule: atoms carry features (a one-hot encoding of atom types plus formal charges), bonds are encoded in a matrix (existence, aromaticity, bond order channels), and coordinates give the 3D conformation. Gaussian noise is progressively added to the coordinates during a forward diffusion process, and a network is trained to predict that noise using mean-squared error.
The distinctive part is where the conditioning comes from. The team uses MoLLaMA, a 960M-parameter molecular foundation model derived from LLaMA 2 and pretrained autoregressively on 1.8B molecules (90B SELFIES tokens) from ZINC-15. They add a set of randomly initialized, learnable query tokens and have them attend to the SELFIES input inside the frozen MoLLaMA, producing a compact bottleneck of chemical information. A trainable projector then processes these query representations together with the SELFIES representations using bi-directional attention, and a feed-forward network maps the output to the dimensionality the diffusion model expects. This condition is added element-wise to the diffusion model's native conditioning (timestep, and optionally a molecular property embedding) and injected via adaLN.
Training happens in two stages. Stage 1 produces high-quality 1D molecules matching the target distribution. Stage 2 freezes MoLLaMA and trains the projector, FFN, and diffusion model end-to-end on the diffusion objective, so the 1D latents are adapted for 3D geometry without fine-tuning the large backbone. The diffusion backbone combines components from the Diffusion Graph Transformer and Diffusion Molecule Transformer, using Relational Multi-head Self-attention; it has 10 layers, 8 heads, and 55M total parameters, and is optimized with AdamW. All experiments run on 4 NVIDIA A800-40GB GPUs.
Why This Matters
Impact on research. The paper argues that prior 3D generation work either models a single modality or fuses modalities heuristically, leaving 1D and 3D generation in isolation. MolSculpt demonstrates that a frozen 1D foundation model can act as a feature resampler providing high-quality conditions for 3D generation, and the ablation showing that a fine-tuned backbone performs worse than a frozen one (FCD_3D 0.969 vs 0.952) is a concrete data point for how to transfer large pretrained molecular models rather than adapt them.
Real-world applications (as framed by the paper):
- Rapid conformer ensemble generation for virtual screening in drug discovery.
- Pocket-conditioned ligand design, where 3D shape and pose matter for binding.
- Diffusion-based docking for scalable, accurate pose prediction.
- Property-conditional molecule design for materials and catalysis, using targets such as HOMO-LUMO gap, dipole moment, polarizability, and heat capacity.
Industry relevance. The work targets pharmaceutical and materials chemistry pipelines where validity, stability, and controllability of generated structures determine whether a candidate is usable. Because Stage 2 only trains a projector, FFN, and a 55M-parameter diffusion model while keeping the 960M-parameter backbone frozen, the approach is also relevant as a template for cheaply repurposing large pretrained scientific models. The code is released at https://github.com/SakuraTroyChen/MolSculpt.
Future Directions
- Scaling to larger biomolecules. The authors explicitly name this as future work, noting that the current evaluation is limited to GEOM-DRUGS (up to 181 atoms, average 44.4) and QM9 (up to 9 heavy atoms, up to 29 including hydrogen).
- Integrating more diverse chemical modalities. The conclusion proposes incorporating additional chemical modalities beyond the 1D SELFIES foundation model to further advance molecule generation.
- Closing remaining geometric gaps. The QM9 FCD_3D result (0.952) trails NExT-Mol (0.879), and the generated C-H bond length distributions are broader and less peaked than the test set — both point to open questions about bond-length fidelity and 3D distribution matching.
- Extending conditional control beyond QM9's six properties. The conditional evaluation covers μ, α, C_v, ε_HOMO, ε_LUMO, and Δε; whether the reported MAE reductions transfer to other property sets or larger datasets is not addressed in the paper.
Target Audience
This paper is most useful to machine learning researchers working on generative models for scientific data, particularly those interested in diffusion models, multimodal conditioning, and cross-modal transfer from large pretrained models. It is also relevant to computational chemists and cheminformatics practitioners who need valid, stable, and property-controllable 3D conformers for screening or design. Readers without a background in diffusion processes, geometric graph representations, or molecular string formats such as SELFIES will find the method section demanding.
Authors’ abstract
Generating precise 3D molecular geometries is crucial for drug discovery and material science. While prior efforts leverage 1D representations like SELFIES to ensure molecular validity, they fail to fully exploit the rich chemical knowledge entangled within 1D models, leading to a disconnect between 1D syntactic generation and 3D geometric realization. To bridge this gap, we propose MolSculpt, a novel framework that "sculpts" 3D molecular geometries from chemical syntax. MolSculpt is built upon a frozen 1D molecular foundation model and a 3D molecular diffusion model. We introduce a set of learnable queries to extract inherent chemical knowledge from the foundation model, and a trainable projector then injects this cross-modal information into the conditioning space of the diffusion model to guide the 3D geometry generation. In this way, our model deeply integrates 1D latent chemical knowledge into the 3D generation process through end-to-end optimization. Experiments demonstrate that MolSculpt achieves state-of-the-art (SOTA) performance in \textit{de novo} 3D molecule generation and conditional 3D molecule generation, showing superior 3D fidelity and stability on both the GEOM-DRUGS and QM9 datasets. Code is available at https://github.com/SakuraTroyChen/MolSculpt.