Skip to content
AI.info

Research

ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry

Overview Research area: Evaluation of multimodal large language models (MLLMs) on chemistry problems that require combining visual, textual, and symbolic information. Technical level: Intermediate. Th

arXiv
2511.17909
Published
2025-11-22
Authors
Zhiyuan Huang, Baichuan Yang, Zikun He, Yanhong Wu, Fang Hongyu, Zhenhe Liu, Lin Dongsheng, Bing Su

AI summary

Overview

Research area: Evaluation of multimodal large language models (MLLMs) on chemistry problems that require combining visual, textual, and symbolic information.

Technical level: Intermediate. The paper is a benchmark-and-evaluation study; it assumes familiarity with multimodal model evaluation and basic chemistry notation (SMILES, structural formulas, crystal diagrams).

Scope: The paper introduces ChemVTS-Bench, a 445-question benchmark built from high-school chemistry exam items, and uses it to measure how six state-of-the-art MLLMs reason when the same problem is posed as an image, as text, or as a SMILES-based symbolic input.

What This Paper Is About

Existing chemistry benchmarks for multimodal models mostly use simple image–text pairs whose images carry weak chemical semantics, which can overstate how well models actually reason about chemistry. The authors build a benchmark where every question is tied to genuinely chemical visual content (molecular structures, inorganic materials, 3D crystal structures) and is presented in three complementary input modes so that modality-dependent behavior can be isolated. The goal is to determine whether current MLLMs can integrate chemically meaningful information across vision, text, and symbolic representations, and to diagnose where they fail.

Key Contributions

  1. A domain-authentic, challenging benchmark. ChemVTS-Bench contains 445 questions drawn from a large corpus of high-school chemistry exam items, filtered first by the Qwen3-VL model and then re-selected by professional chemistry teachers. Questions are grouped into three categories — structural chemistry, organic chemistry, and others — and include multiple-choice and fill-in-the-blank formats. Each item ships with a standard answer and a reference explanation, plus fine-grained annotations of the visual skills and chemical knowledge points required.
  2. Three complementary input modes per problem. Each task is presented as (1) visual-only, (2) visual–text hybrid, and (3) SMILES-based symbolic input. Due to limits in the expressiveness of SMILES, 95 questions ultimately contain both the visual and the SMILES modality.
  3. A systematic, agent-based evaluation framework. A two-stage automated workflow standardizes inference (consistent prompt templates across modalities, temperature set to 0) and then scores answers and diagnoses failure modes with a defined error taxonomy.
  4. Extensive experimental analysis. Six representative MLLMs — GPT-5, Gemini-2.5-Pro, and Doubao 1.6 Vision (closed-source) and Qwen3-VL, InternVL 3, and Llama 3.2 (open-source) — are compared across domains and modalities. The authors state that all data and code will be released.

Main Findings

  • Visual-only input remains the hardest setting. Average accuracies on the visual-only mode are GPT-5 66.29, Gemini-2.5-Pro 75.73, Doubao 1.6 Vision 65.17, Qwen3-VL 66.07, InternVL 3 34.38, and Llama 3.2 33.71.
  • Text-only input is generally the easiest. The paper reports Gemini-2.5-Pro and InternVL 3 as the best and worst performers in this setting, with average accuracies of 85.26 and 51.58 respectively; GPT-5 scores 81.05, Doubao 1.6 Vision 76.84, Qwen3-VL 69.47, and Llama 3.2 57.89.
  • Visual–text fusion does not restore visual-only performance to text-only levels. Reported visual–text averages are Gemini-2.5-Pro 74.85, Liu Doubao 1.6 Vision 65.03, Qwen3-VL 60.74, GPT-5 57.67, Llama 3.2 47.24, and InternVL 3 34.97. The paper states that in the visual–text setting all models showed an increase in average accuracy and that vision-related errors persisted or even increased; the per-model averages in Table 2 show mixed directions relative to each model's visual-only score.
  • Structural chemistry is the hardest domain. The authors attribute this to visual–symbolic complexity: questions involve crystal lattices, organic frameworks, or novel compound structures not encountered in pretraining, and many diagrams are not easily captured by OCR, leading to misinterpretation or hallucinated structural features.
  • Organic chemistry is more approachable. When problems concern reaction types, functional groups, or substitution rules, models often answer from textual reasoning alone, and SMILES representations let them parse full molecular structures. Failures cluster on newly synthesized or uncommon compounds.
  • The "others" category scores highest. These questions mostly test chemical common sense or fundamental principles such as periodic trends. For element inference tasks, the paper notes most elements involved are main-group atoms with atomic number below 30, whose properties are well covered in pretraining corpora.
  • Cross-modal consistency varies sharply. Using consistency (the proportion of instances where a model is correct on all modalities or incorrect on all), Gemini-2.5-Pro leads across modality groups (VTS 86.67, V&T 90.53, VT&V 84.66, VT&T 90.00). Open-source models align better on VT&T (Qwen3-VL 86.67, InternVL 3 83.33) but degrade on settings involving visual-only input. Llama 3.2 shows the largest discrepancies (VTS 50.00, V&T 66.32, VT&V 61.96, VT&T 80).
  • Visual recognition errors dominate in unimodal settings. Across all models, VRE is the most prevalent error type for visual-only tasks; symbolic parsing and numerical computation errors stay consistently low across modalities, which the authors read as evidence that current architectures handle chemical text and mathematical expressions adequately.
  • Knowledge-based and logical reasoning errors persist across modalities. The paper argues these stem from training corpus and architecture rather than input format.
  • The error-diagnosis agent matches human judgment. On a set of 30 representative error instances generated by various MLLMs, the agent achieved a 100% consistency rate with manual verification.
  • Coverage is validated by an independent teacher panel. Chemistry teachers who did not take part in question screening analyzed full sets of comprehensive high-school chemistry exam papers; the authors report strong coverage of the major content areas of the high-school chemistry curriculum.
  • Prior benchmarks are comparatively narrow. The paper's comparison table shows ChemIQ (no vision, SMILES yes, reasoning yes, structure representation no), ChemTable (vision yes, no SMILES, reasoning yes, no structure representation), ChemBench (no vision, SMILES yes, no reasoning, no structure representation), and MMCR-Bench (vision yes, no SMILES, reasoning yes, no structure representation), against ChemVTS (all four yes).

Methodology in Plain English

The authors started from a large bank of high-school chemistry exam questions and kept only items whose images are essential to solving the problem. A vision-language model (Qwen3-VL) did a first pass to find image-containing questions, then human chemistry teachers did a second pass, judging whether each image carries real chemical content and whether the problem requires genuine reasoning at moderate difficulty.

Each retained question was split into a text component and a visual component, with chemical structures segmented into separate images; LaTeX was kept for equations and formulas, and special tokens marked where images sit inside the text. Questions were then categorized as structural chemistry, organic chemistry, or others. Where possible, structures were converted into SMILES strings so the same problem could also be posed symbolically.

For evaluation, the same question is fed to a model in three modes: image only, text only, and image plus text. Prompt templates share one system prompt and differ only in the user prompt, so differences in scores trace back to the input mode rather than to phrasing. All inference used a temperature of 0.

Scoring is automated. A Doubao-based evaluation agent compares the model's output against the ground-truth solution, allowing permissible variation in expression and intermediate steps, and returns a binary correct/incorrect decision. For responses judged incorrect, a second Doubao-based analysis agent assigns the failure to one of five predefined categories: Visual Recognition Error, Symbolic Parsing Error, Numerical Computation Error, Knowledge-based Error, or Logical Reasoning Error. Each category has tailored prompts — for example, the VRE prompt triggers on phrases such as "in the figure," "in the image," or "as shown in the figure" combined with verbs like "misidentified as," "mistaken for," "overlooked," or "misinterpreted."

Why This Matters

Impact on research. The benchmark separates modality effects from task difficulty, which earlier chemistry benchmarks could not do. Because the same question exists in visual-only, visual–text, and SMILES form, researchers can tell whether a model fails because it cannot see, because it cannot parse symbols, or because it lacks chemistry knowledge. The finding that visual grounding — not symbolic parsing or arithmetic — is the persistent bottleneck gives the field a concrete target for improvement.

Real-world applications (implications the benchmark's design speaks to):

  • Automated or assisted grading of chemistry exam items that contain structures, diagrams, and multi-part questions.
  • Software that reads molecular diagrams, crystal structures, and reaction schematics directly from figures rather than from typed formulas.
  • Materials and crystal-structure workflows where spatial reasoning over 3D lattices matters.
  • Drug and molecule discovery pipelines in which models must move between drawn structures, SMILES strings, and natural-language descriptions.

Industry relevance. The evaluation uses an automated agent workflow rather than manual grading, which is what makes large-scale, reproducible chemistry benchmarks practical for labs and companies. The paper's comparison with prior benchmarks also warns that evaluations built on images with weak chemical semantics can overestimate model capability, a risk for anyone selecting a model for chemistry work.

Future Directions

  • Close the visual grounding gap. Since visual recognition errors dominate in unimodal settings, the paper's results point to better chemical image understanding as the highest-value improvement area, especially for crystal structures and molecular diagrams that OCR cannot capture.
  • Address knowledge-based and logical reasoning errors. These persist across all modalities, so they cannot be fixed by adding text or symbolic inputs; the authors attribute them to training corpus and architecture.
  • Extend symbolic expressiveness. Only 95 questions currently carry both visual and SMILES modalities because of limits in SMILES expressiveness; finding representations for structures that cannot be linearized is an open problem.
  • Broaden model and domain coverage. The present study covers six models and three chemistry categories drawn from high-school curriculum content; whether the same modality-dependent patterns hold for larger, expert-level chemical problems is not reported.

Target Audience

Researchers and engineers working on multimodal large language models, particularly those building or evaluating models for scientific and chemical domains. It is also useful for chemistry education researchers and assessment designers interested in automated grading of structure-heavy questions, and for practitioners who need to know how much a general-purpose MLLM can be trusted with chemical diagrams before deploying it.

Authors’ abstract

Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs with limited chemical semantics. As a result, the actual ability of Multimodal Large Language Models (MLLMs) to process and integrate chemically meaningful information across modalities remains unclear. We introduce \textbf{ChemVTS-Bench}, a domain-authentic benchmark designed to systematically evaluate the Visual-Textual-Symbolic (VTS) reasoning abilities of MLLMs. ChemVTS-Bench contains diverse and challenging chemical problems spanning organic molecules, inorganic materials, and 3D crystal structures, with each task presented in three complementary input modes: (1) visual-only, (2) visual-text hybrid, and (3) SMILES-based symbolic input. This design enables fine-grained analysis of modality-dependent reasoning behaviors and cross-modal integration. To ensure rigorous and reproducible evaluation, we further develop an automated agent-based workflow that standardizes inference, verifies answers, and diagnoses failure modes. Extensive experiments on state-of-the-art MLLMs reveal that visual-only inputs remain challenging, structural chemistry is the hardest domain, and multimodal fusion mitigates but does not eliminate visual, knowledge-based, or logical errors, highlighting ChemVTS-Bench as a rigorous, domain-faithful testbed for advancing multimodal chemical reasoning. All data and code will be released to support future research.

Read the original paper