Skip to content
AI.info

Research

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Overview Research area: Computer Vision, specifically Mathematical Expression Recognition (MER) within document parsing and understanding, with comparisons against general-purpose multimodal large lan

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline
arXiv
2512.13731
Published
2025-12-14
Authors
Weikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie, Zhineng Chen

AI summary

Overview

  • Research area: Computer Vision, specifically Mathematical Expression Recognition (MER) within document parsing and understanding, with comparisons against general-purpose multimodal large language models (MLLMs).
  • Technical level: Advanced. The paper assumes familiarity with encoder-decoder architectures, tokenizers, autoregressive decoding, LaTeX, syntax trees, and MER evaluation metrics.
  • Scope: The paper builds a difficulty-stratified benchmark (CMER-Bench), two large training datasets (MER-17M and CMER-3M), a specialized tokenizer and target representation (Structured Mathematical Language), and a 125-million-parameter baseline model (CMERNet) for recognizing complex, multi-line mathematical expressions.

What This Paper Is About

Mathematical Expression Recognition has advanced on simple, single-line formulas, but models still fail on complex expressions that contain hundreds of tokens and span multiple lines. The authors argue this failure stems from three gaps: no benchmark that measures difficult expressions, training data that is small and heavily skewed toward short/simple samples, and a representation dilemma in which linear LaTeX strings are easy to generate but hard to align with two-dimensional layouts, while syntax-tree formats are intuitive but awkward for today's autoregressive decoders. The goal is to provide a benchmark, large-scale data, a structure-aware representation, and a specialized model that together close this gap.

Key Contributions

  1. CMER-Bench, a curated benchmark of 2,000 samples split into Easy, Moderate, and Complex difficulty tiers, built through a hybrid process combining automated difficulty scoring with human expert verification. The authors state it is the first to cover evaluation of complex expressions.
  2. MER-17M and CMER-3M, two new training datasets. MER-17M contains over 17 million expression image-LaTeX pairs (17.7M in Table 1) and CMER-3M is a length-balanced subset of over 3 million samples (3.1M in Table 1). The authors describe CMER-3M as the most complex and MER-17M as the largest MER training dataset to their knowledge.
  3. A specialized BPE tokenizer and Structured Mathematical Language (SML), a new target representation that parses LaTeX into a syntax tree and then serializes that tree into a linear token sequence carrying explicit structural grammar tokens, retaining compatibility with autoregressive decoders.
  4. CMERNet, a 125-million-parameter encoder-decoder baseline with a CNN-Transformer hybrid vision encoder, a feature-refining connector, and a cross-attention-based autoregressive decoder, trained end-to-end on CMER-3M.

Main Findings

  • Existing models collapse on complex expressions: On the Complex tier of CMER-Bench, BLEU scores fall to 0.3089 for Gemini-2.5-pro, 0.2082 for ChatGPT-4o, and 0.1915 for Qwen-VL-72b. The paper reports that all compared methods, especially general-purpose MLLMs, show a drastic performance drop relative to CMERNet.
  • Dedicated models fare only slightly better: UniMERNet (a dedicated MER method) and the document parsing models MinerUv2 and Dolphin obtain slightly better results than general MLLMs, which the authors attribute to their training on expression data.
  • CMERNet leads on most metrics across all three tiers: CMERNet records the best ROUGE-1, ROUGE-2, ROUGE-L, BLEU, and lowest Average Edit Distance in every difficulty tier. Examples: Easy tier ROUGE-1 0.8626, BLEU 0.7653, Avg Edit Dist. 51.61; Moderate tier ROUGE-1 0.8266, BLEU 0.7215, Avg Edit Dist. 88.09; Complex tier ROUGE-1 0.7750, BLEU 0.5568, Avg Edit Dist. 474.18.
  • The BLEU advantage is reported as 15%-29%: The authors state CMERNet excels the second-best method in Table 4 by nearly 15%-29% in BLEU score under the different difficulty tiers.
  • CDM does not show the same gaps: On the CDM metric, CMERNet scores 0.966 (Easy), 0.845 (Moderate), and 0.690 (Complex); Dolphin scores 0.886 on Moderate and Gemini-2.5-pro scores 0.856 on Complex, so CDM is not universally led by CMERNet. The paper explains this is because CDM mainly measures bounding-box-level matching and hardly perceives subtle mistakes.
  • Length and layout distributions are heavily skewed in prior data: In Table 1, UniMER contains 1.1M samples while CMER-3M contains 3.1M and MER-17M contains 17.7M. Prior datasets concentrate on short expressions (IM2LATEX 75.3K total, Pix2tex 233.8K total), whereas CMER-3M has 858.1K samples in the 151-300 token range and 855.6K in the 301-450 range.
  • CMER-3M is structurally more complex: Table 2 reports CMER-3M has more than 2.5M multi-line expressions, accounting for 83% of its samples, a ratio the paper says significantly exceeds existing datasets.
  • Dynamic resolution resizing causes negligible distortion: With a patch size of 16 on 1024x1024 input, the theoretical maximum distortion ratio is under 2.93%; with patch size 32 it is under 6.05% (Table 3).
  • Training cost and setup: The model was trained for one epoch using AdamW (β1 = 0.9, β2 = 0.999), a linear warmup from 1×10⁻⁵ to 1×10⁻⁴ over the first 3,000 steps, then cosine decay to 1×10⁻⁹, on eight NVIDIA RTX 4090 GPUs with Distributed Data Parallel, a per-GPU batch size of 8 and a global batch size of 64; the whole training took around 3 days.

Methodology in Plain English

The authors start by collecting more than 1 million scientific documents from repositories such as ArXiv, extracting LaTeX snippets with parsing tools, removing visually/semantically misaligned tags such as cite, rendering the expressions into images, and filtering out malformed or overly simplistic examples while applying diversity checks. This yields over 17 million image-LaTeX pairs; 2,000 of them become the CMER-Bench evaluation set, and the rest form MER-17M. Because MER-17M's distribution is unbalanced, they build CMER-3M as a more length-balanced subset.

For the model, they replace the general-purpose tokenizer with a custom Byte Pair Encoding tokenizer. They first designate common LaTeX environments and commands such as sqrt, frac, and begin{gather} as special tokens so they are treated as atomic units, then train BPE on CMER-3M. For the output format, they parse each LaTeX string into a syntax tree and then serialize that tree into a token sequence that encodes parent-child relationships, node types, and sibling order through additional grammar tokens. This is the Structured Mathematical Language representation, which keeps a linear output shape but adds explicit structure.

Images are preprocessed with CMER-Fit, a parameter-free dynamic resolution module inspired by NaViT and SigLip2, which scales each image to the largest size whose patch grid fits within a fixed token budget (for example, 256 patches), avoiding distortion from fixed-size resizing. The model itself uses six stacked residual CNN blocks without downsampling in early layers for fine-grained detail, followed by 12 standard Transformer layers for global context, two MLP layers to align visual features with linguistic representations, and a cross-attention Transformer decoder that autoregressively produces the SML sequence. Training mixes CMER-3M with CROHME2013, CROHME2016, CROHME2019, CROHME2023, HME-100K, and UniMER-1M.

Why This Matters

Complex mathematical expression recognition sits underneath scientific document parsing, accessibility tools, and technical search. By showing that strong MLLMs lose most of their accuracy when expressions become long and multi-line, the paper identifies a concrete blind spot in current document AI evaluation, and it provides the data and benchmark needed to measure progress on it. Its main conceptual point is representational: how you encode a formula as a training target matters, and structure-aware linear targets may be a better fit for today's autoregressive models than plain LaTeX.

Real-world applications include:

  • Scientific document digitization: Converting PDFs and scanned papers into machine-readable, editable formulas for publishers, archives, and literature databases.
  • Accessibility: Producing accurate spoken or Braille renderings of multi-line equations for visually impaired readers.
  • Automated grading and assessment: Parsing students' handwritten or typeset mathematics in education platforms, including multi-line derivations.
  • Technical search and retrieval: Indexing formulas so search engines and knowledge bases can match expressions structurally rather than as plain strings.

Industry relevance centers on any product pipeline that ingests technical documents: academic publishing, patent analysis, scientific knowledge bases, education technology, and document AI platforms. The paper's release of MER-17M and CMER-3M as public training corpora also lowers the data barrier for groups that lack the resources to scrape and clean millions of scientific expressions themselves.

Future Directions

  • Extending CMERNet's validation: The authors explicitly state they plan to conduct more experiments to further verify CMERNet.
  • Strengthening MLLMs with the new data: The paper proposes exploring the use of the constructed datasets to improve the MER capability of popular MLLMs, rather than only training a specialized model.
  • Closing the CDM gap: Because CDM relies on bounding-box-level matching and does not reveal the subtle errors that BLEU and ROUGE expose on complex expressions, better evaluation metrics for fine-grained structural fidelity remain an open problem.
  • Scaling the architecture and representation: The paper does not report whether SML and the hybrid encoder continue to improve at larger parameter counts or with interleaved CNN-Transformer designs, leaving the scaling behavior of this approach untested.

Target Audience

This paper suits researchers and engineers working on document understanding, OCR, formula recognition, and multimodal LLMs, particularly those building or evaluating systems that must handle scientific and mathematical content. It is also relevant to dataset builders and benchmark designers interested in difficulty stratification and data-skew analysis, and to practitioners in publishing, education technology, or accessibility who need to assess whether current models are trustworthy on complex equations. Readers should already be comfortable with encoder-decoder architectures, tokenization, and standard sequence-generation metrics such as BLEU, ROUGE, and edit distance.

Authors’ abstract

Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench.

Read the original paper