Research
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
Overview Research area: Multimodal vision-language modeling for chemistry, combining computer vision, molecular representation learning, and 3D geometric descriptors (drug discovery, optical chemical

- arXiv
- 2601.14732
- Published
- 2026-01-21
- Authors
- Jing Lan, Hexiao Ding, Hongzhao Chen, Yufeng Jiang, Nga-Chun Ng, Gwing Kei Yip, Gerald W. Y. Cheng, Yunlin Mao, Jing Cai, Liang-ting Lin, Jung Sun Yoo
AI summary
Overview
Research area: Multimodal vision-language modeling for chemistry, combining computer vision, molecular representation learning, and 3D geometric descriptors (drug discovery, optical chemical structure recognition, molecular captioning, molecular property prediction).
Technical level: Advanced. The paper specifies full architectural equations, tokenization algorithms, hashing schemes, and two-stage training hyperparameters.
Scope: The paper proposes DeepMoLM, a dual-view framework that grounds 1024 × 1024 molecular images in 3D conformer-derived geometric fingerprints, and evaluates it on three tasks: PubChem molecule captioning, computed property prediction, and ChEBI-20 molecule description from images.
What This Paper Is About
Most molecular language models read molecules as 1D strings (SMILES/SELFIES) or 2D graphs, while vision-language models either lose stereochemical detail or fail to connect continuous 3D structure to discrete language tokens. The authors argue that rotation/translation of the same molecule can change its embedding, that enantiomers and stereoisomers can become indistinguishable, and that high-frequency visual cues such as stereobonds and ring closures are lost at standard image resolutions. DeepMoLM's goal is to fuse high-resolution molecular images with explicit 3D structural descriptors so the model can generate physically grounded molecule-to-text output without needing atom coordinates at inference.
Key Contributions
- A dual-pathway Molecular DeepEncoder that processes 1024 × 1024 molecular images and preserves fine-grained stereochemical markers through convolutional token compression, keeping the token budget manageable despite the high input resolution.
- A cross-attention fusion projector that aligns visual token representations with discrete 3D Extended 3-Dimensional Fingerprint (E3FP) tokens, providing explicit geometric grounding inside a unified multimodal embedding space rather than simple concatenation.
- An end-to-end vision-language pipeline built on Qwen2-VL that performs 3D molecule captioning and property prediction from images alone, without requiring atom coordinates at inference.
- Empirical evidence that this geometric grounding improves molecule captioning and property prediction, both of which are tied to stereochemistry.
Main Findings
- PubChem molecule captioning gains: DeepMoLM's generalist variant scored BLEU-2 35.87, BLEU-4 25.53, ROUGE-1 47.07, ROUGE-2 23.07, ROUGE-L 30.76, and METEOR 37.66, a 12.3% relative METEOR gain over the strongest generalist baseline. It exceeded generalist baselines by more than 10% relative gain in METEOR; ROUGE-1 44.49 and METEOR 35.87 in the specialist setting, where it matched UniMoT (ROUGE-1 37.50, METEOR 34.80).
- Full validity on property queries: DeepMoLM produced valid numeric outputs for all property queries, reaching a 100% valid answer rate across Molecular Weight, LogP, TPSA, and Complexity, avoiding the formatting failures seen in generalist models.
- Specialist property-prediction accuracy: MAE 13.64 g/mol on Molecular Weight (best, vs 3D-MoLM at 14.79 and Uni-Mol at 20.35) and 37.89 on Complexity (best, vs 3D-MoLM at 44.85), while remaining competitive on LogP (0.73) and TPSA (10.04).
- Generalist property-prediction accuracy: MAE 14.63 on Molecular Weight and 41.73 on Complexity with 100% valid answers, outperforming 3D-MoLM (16.58 and 45.49) and Qwen2-VL-7B (42.13 and 103.07).
- ChEBI-20 description from images: DeepMoLM reached BLEU-2 55.76, BLEU-4 46.29, ROUGE-2 48.07, ROUGE-L 55.96, and METEOR 58.23, exceeding generalist baselines and matching state-of-the-art vision-language models such as Mol-VL-7B (55.73 / 46.14 / 47.26 / 56.61 / 58.14). Specialist string-based models such as BioT5+ (66.6 / 59.1 / 58.4 / 65.0 / 68.1) still scored higher.
- Ablation results: Removing pre-training caused a clear drop across all metrics. Replacing the fusion projector with concatenation performed worse than the plain non-pretrained model on description, with the largest losses on BLEU-2 (19.94) and METEOR (27.29). Removing the 3D-E3FP branch and using a linear projection yielded the weakest ROUGE-L, indicating degraded semantic fidelity.
- Case study discrepancy: In the reported case study for an open-text drug description, DeepMoLM's output described a branched amino tetrasaccharide, while the ground truth described a branched amino trisaccharide consisting of N-acetyl-beta-D-glucosamine with beta-D-glucuronosyl and N-acetyl-beta-D-glucosaminyl residues at the 3- and 6-positions.
Methodology in Plain English
The authors treat molecule-to-text generation as a multimodal autoregressive task. Each molecule is given to the model twice: once as a 1024 × 1024 image, and once as a structural token sequence built from canonical SELFIES tokens combined with E3FP identifiers.
For the image, a three-part encoder is used. A 12-layer SAM-Base local vision transformer processes 16-pixel patches (4096 tokens) with window attention of size 14 to capture fine detail. A convolutional token compressor — two 3 × 3 convolutions with stride 2 and padding 1 — shrinks the 64 × 64 token grid to 16 × 16, i.e. 256 tokens, keeping the channel dimension at 1024. A 24-layer CLIP-Large global transformer then applies dense global attention to those 256 tokens. Global and local tokens are concatenated channel-wise into a 256 × 2048 representation.
For the geometry, the heavy-atom conformer coordinates are converted into E3FP fingerprints: an atom-level invariant is hashed with MurmurHash3, then for K iterations, neighbors within radius R_j = r · j are collected along with connectivity and stereochemical configuration, sorted, and re-hashed. Each hashed identifier is mapped into a discrete vocabulary of size |F| by taking it modulo |F|. SELFIES and E3FP tokens are aligned through a heavy-atom index mapping, and their embeddings are averaged (1D embedding plus a masked 3D embedding, divided by 1 + m_t) into a fused structural sequence.
A post-normalization Transformer block with multi-head cross-attention then fuses the two streams, using visual tokens as queries and structural tokens as keys and values, with padding masked out and no causal mask since fusion is not autoregressive. The output keeps 256 visual tokens and projects into the decoder's hidden size (4096 for the 7B decoder). A Qwen2-VL decoder then concatenates the fused visual tokens with the embedded text prompt and generates captions and property text.
Training is two-stage. In Stage-1, the encoder and decoder are frozen and only the fusion projector is trained on molecular image-text pairs. In Stage-2, the encoder stays frozen while the fusion projector and decoder are jointly fine-tuned on instruction-style data. Everything is trained on a single node of 8 NVIDIA H800 GPUs (80GB PCIe) in PyTorch, using AdamW, BF16, DeepSpeed ZeRO-3, 5 pre-training epochs at learning rate 1e-4 (batch 128, warmup 0.03, context 4096) and 10 fine-tuning epochs at up to 5e-5 (batch 64, warmup 0.01). Two variants are reported: a Generalist model trained on a mixture of all downstream datasets, and a Specialist model fine-tuned separately per task. The decoder backbone is Qwen2-VL-7B-Instruct and the image encoder is a pretrained SAM-CLIP Molecular DeepEncoder.
Why This Matters
The work targets a concrete gap between how chemical knowledge is stored (images in papers and patents) and how models interpret it (strings and 2D graphs). By linking images to 3D-aware geometric invariants and removing the need for atom coordinates at inference, it points toward models that can mine chemical literature directly from figures while respecting stereochemistry — the property that determines whether a molecule is a drug or its inactive mirror image.
Real-world applications:
- Automated chemical literature and patent mining, extracting structured, stereochemistry-aware descriptions from molecular figures in PDFs and scanned documents.
- Drug discovery pipelines, where molecular weight, LogP, TPSA, and complexity are routinely used for filtering and triage, and where the model's 100% valid-output rate matters for automated downstream processing.
- Chemical database curation and digitization, converting legacy image-only compound records into captions and property values.
- Educational and QA interfaces for chemistry, as illustrated by the open-text molecular QA interface in Figure 3 of the paper.
Industry relevance: the paper explicitly frames the problem around drug discovery and chemical literature mining. Because the model matches specialist string-based systems without using SMILES at inference, it is relevant to settings where structures exist only as images, or where visual rendering is more reliable than textual encoding. The released code (https://github.com/1anj/DeepMoLM) makes the approach reproducible.
Future Directions
- Closing the specialist gap on description generation. BioT5+ (BLEU-2 66.6, METEOR 68.1) still outperforms DeepMoLM (55.76, 58.23) on ChEBI-20, so large-scale 1D sequence pre-training remains an advantage the authors only partially narrow.
- Reducing naming and structural errors in open-text generation. The reported case study mismatch (tetrasaccharide output vs. trisaccharide ground truth) indicates room for improvement in complex oligosaccharide descriptions.
- Extending beyond the four properties evaluated. Only Molecular Weight, LogP, TPSA, and Complexity were tested; whether the geometric grounding transfers to stereochemistry-dependent or quantum-level properties is not established here.
- Scaling the resolution/accuracy trade-off further. The convolutional token compression is what makes 1024 × 1024 inputs affordable; whether higher input resolutions yield further gains is presented as a cost trade-off (quadratic self-attention) rather than settled.
Target Audience
Researchers and practitioners working on multimodal molecular machine learning, vision-language models for chemistry, and optical chemical structure recognition, as well as cheminformatics and drug-discovery teams that need to extract structured knowledge from molecular images. Readers should be comfortable with Transformer architectures, cross-attention, and molecular representations such as SELFIES and fingerprints; the ablation and benchmark tables are also directly useful to those comparing against MolT5, 3D-MoLM, UniMoT, Mol-VL, and BioT5+ baselines.
Authors’ abstract
AI models for drug discovery and chemical literature mining must interpret molecular images and generate outputs consistent with 3D geometry and stereochemistry. Most molecular language models rely on strings or graphs, while vision-language models often miss stereochemical details and struggle to map continuous 3D structures into discrete tokens. We propose DeepMoLM: Deep Molecular Language M odeling, a dual-view framework that grounds high-resolution molecular images in geometric invariants derived from molecular conformations. DeepMoLM preserves high-frequency evidence from 1024 $\times$ 1024 inputs, encodes conformer neighborhoods as discrete Extended 3-Dimensional Fingerprints, and fuses visual and geometric streams with cross-attention, enabling physically grounded generation without atom coordinates. DeepMoLM improves PubChem captioning with a 12.3% relative METEOR gain over the strongest generalist baseline while staying competitive with specialist methods. It produces valid numeric outputs for all property queries and attains MAE 13.64 g/mol on Molecular Weight and 37.89 on Complexity in the specialist setting. On ChEBI-20 description generation from images, it exceeds generalist baselines and matches state-of-the-art vision-language models. Code is available at https://github.com/1anj/DeepMoLM.