Skip to content
AI.info

Research

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents

Overview Research area: Construction and serving of large-scale structured organic reaction data for chemistry research and AI for chemistry (AI4Chem), combining document information extraction, chemi

DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents
arXiv
2609.06703
Published
2026-09-06
Authors
Yubin Wang, Xingjian Wei, Jiang Wu, Yinfan Wang, Boyu Zhu, Lin Zhang, Jianing Yu, Huazheng Zeng, Ruiyi Ding, Junyuan Gao, Jiaxing Sun, Lingli Ge, Haote Yang, Jingchao Wang, Aijia Guo, Qian Jiang, Yurui Zhao, Wenjian Zhang, Chen Zhu, Lijun Wu, Xiaolei Yang, Haodong Chen, Junjie Yuan, Zichao Ye, Shaowei Hou, Jing Ye, Jia Yu, Shan Wang, Lijun Wu, Jiantao Qiu, Chao Xu, Yuqiang Li, Guangyu Wang, Bowen Zhou, Dahua Lin, Conghui He

AI summary

Overview

Research area: Construction and serving of large-scale structured organic reaction data for chemistry research and AI for chemistry (AI4Chem), combining document information extraction, chemical data normalization, and retrieval infrastructure.

Technical level: Intermediate. The report is deliberately written at "capability level," presenting data sources, object design, evaluation protocols, and system interfaces; the authors state that the specific technical details of the extraction pipeline will be described in a subsequent technical report.

Scope: DianShi-RxnDB is a large-scale, fine-grained organic reaction data platform built from USPTO and EPO patent documents by a fully automated extraction and normalization pipeline, served through both a Web research workbench for researchers and an MCP service for AI agents.

What This Paper Is About

Useful reaction records need to capture not only the structural change between reactants and products but also participant roles, quantities, conditions, yield, and experimental procedure, and much of this knowledge is dispersed across patent text, images, tables, and reaction schemes. Scaling the extraction of such fine-grained, provenance-linked records to large document collections, and making them accessible both to human researchers and to AI agents, remains a central challenge.

The paper introduces DianShi-RxnDB, which transforms approximately 1.58 million USPTO and EPO patent documents (primarily covering 1976–2025) into approximately 24 million single-step Reaction Instances, of which approximately 14.8 million (61.7%) pass an automated qualification assessment, and exposes them through a Web research workbench and a Model Context Protocol (MCP) service.

Key Contributions

  1. A large-scale, fine-grained reaction data foundation built by a fully automated pipeline. A fully automated information-extraction and normalization pipeline covering patent text, images, and reaction schemes produces approximately 24 million Reaction Instances from USPTO and EPO organic synthesis patents, of which approximately 14.8 million pass automated qualification. The authors report a manual quality evaluation of five core fields on a random sample of 1,300 qualified instances, with a micro-averaged accuracy of 92.95%, plus an external matched comparison with the Pistachio Reaction Dataset.

  2. Instance-level organization of reaction knowledge with provenance linkage. Each Reaction Instance records participant roles (Reactant, Product, Reagent, Solvent, Catalyst), quantities and equivalents, conditions such as temperature and time, yield, and experimental procedures, and is linked to its source patent and source location. The database establishes structured relationships among Substance, Reaction Instance, Reaction Group, and Reference objects (and additionally maintains separate Reaction Template collections from LocalRetro and RDChiral).

  3. Complementary researcher and AI-agent interfaces over the same data foundation. The Web research workbench supports search, filtering, instance comparison, linked exploration, and source-patent verification, while the MCP service provides composable structured retrieval tools. The authors demonstrate the two interfaces and their complementary use in the same reaction-precedent retrieval task.

Main Findings

  • Corpus scale and funnel: The initial collection contained 20,256,438 patent records (12,377,492 USPTO and 7,878,946 EPO). IPC-based filtering for organic chemistry relevance retained 2,512,925; within-source consolidation and deduplication left 1,968,230; cross-source deduplication between USPTO and EPO left 1,757,755; removal of records without usable full text left 1,580,939 patent documents entering the extraction pipeline. Of these, 608,309 ultimately link to at least one retained Reaction Instance.

  • Database object counts: 608,309 source patent documents (References), 6,261,797 Substances, 23,999,236 Reaction Instances, 6,580,963 Reaction Groups, 1,190,000 deduplicated LocalRetro Reaction Templates, and 1,860,000 deduplicated RDChiral Reaction Templates.

  • Automated qualification: 14,808,205 Reaction Instances (61.7%) are qualified; 9,191,031 (38.3%) are not qualified. Qualification uses RXNMapper to generate atom mappings and checks whether the mapped reaction passes atom-conservation checks and other implemented rules. Non-qualified instances are still retained with their extracted experimental information and provenance relationships.

  • Manual field-level quality: Across 1,300 randomly sampled qualified instances and five fields, 6,042 of 6,500 field-level judgments were correct, a micro-averaged accuracy of 92.95%. Field-level results: Yield 1,233/1,300 (94.85%), Reactant 1,226/1,300 (94.31%), Reagent 1,100/1,300 (84.62%), Catalyst 1,265/1,300 (97.31%), Solvent 1,218/1,300 (93.69%). Catalyst was highest and Reagent lowest; Yield, Reactant, and Solvent ranged from 93.69% to 94.85%.

  • Recurring error categories: (1) workup and purification information recorded as reagents or solvents; (2) participant duplication and incorrect role assignment among reactant, reagent, solvent, and catalyst; (3) compact expressions and cross-paragraph extraction, such as omitted parenthetical/concentration solvents and missed yields in later paragraphs; (4) contextual references and reaction-step boundaries, such as unresolved references to general procedures or intermediate preparations and multiple reaction steps combined into one record.

  • Scale comparison with Pistachio: DianShi-RxnDB contains approximately 24 million Reaction Instances versus approximately 21 million reported Pistachio reaction records; the authors note these resource-level counts are not strictly commensurate. In a matched sample of 100 US patents under a common reactant–product-based deduplication rule, DianShi-RxnDB went from 4,222 to 4,093 retained records (129 duplicates removed) and Pistachio (text + image) went from 3,630 to 2,992 retained records (638 duplicates removed: 115 text-internal, 89 image-internal, and 434 text–image overlaps). DianShi-RxnDB thus contained 1.368 times as many records retained after deduplication in this sample, a ratio the authors describe as descriptive of the matched sample and not extrapolated to the complete corpora.

  • Representation granularity: In the inspected shared-reaction example from patent US10975080 (Intermediate 4), DianShi-RxnDB explicitly separates selected participant, process, workup, provenance, and validation features that are not exposed as corresponding standalone fields or objects in the inspected Pistachio record. Specifically, separate Reagent and Catalyst roles, a stage-level reaction object, a dedicated workup field, and explicit validation diagnostics were marked "Yes" for DianShi-RxnDB and "No" for the inspected Pistachio record. Both records contain product structures, quantities, temperature, time, yield information, and an ordered procedure, organized differently.

  • Field-level exact agreement: Across 660 source-paragraph-matched reaction pairs from 58 patents, DianShi-RxnDB showed numerically higher exact agreement against source-grounded references for all six evaluated fields: Reactant 643/660 (97.42%) vs 623/660 (94.39%), gap +3.03%; Reagent 576/660 (87.27%) vs 555/660 (84.09%), +3.18%; Catalyst 569/660 (86.21%) vs 553/660 (83.79%), +2.42%; Solvent 629/660 (95.30%) vs 615/660 (93.18%), +2.12%; Product 609/660 (92.27%) vs 604/660 (91.52%), +0.75%; Yield 656/660 (99.39%) vs 535/660 (81.06%), +18.33%. The authors note the Yield difference is affected in part by representation policy: DianShi-RxnDB generally records explicit percentages from source text, while Pistachio may provide values inferred from mass and stoichiometry when no percentage is explicitly stated. GPT-5.6-sol performed LLM-assisted semantic adjudication, while deterministic code performed normalization, agreement assessment, counting, and percentage calculation.

  • Resource positioning: In the comparison table of representative resources, USPTO-Lowe is listed at approximately 3.75 M (1976–2016), CJHIF at approximately 3.22 M (literature coverage not specified), USPTO-LLM at approximately 247,000 (1976–2016), Pistachio at approximately 21 M (1971–present, paragraph provenance, paid Web access), Reaxys at approximately 73 M (1771–present, document provenance, paid Web access plus MCP/API license), and CAS Reactions at approximately 150 M (1840–present, document provenance, paid Web access). DianShi-RxnDB is listed at approximately 24 M (1976–2025) with paragraph provenance and free non-commercial Web/MCP access.

  • Interfaces: The Web research workbench provides three search entry points (Substance by name or structure, Reaction by Reaction SMILES or Reaction SMARTS returning results at Reaction Instance and Reaction Group granularity, and Reference by patent identifier or text over titles, abstracts, and keywords), linked exploration across object relationships, in-group instance comparison, and precise navigation from a Reaction Instance to the relevant source-patent paragraph or text region. The MCP service provides four categories of structured retrieval: Substance retrieval (by name, database identifier, or chemical representation, including structure-based similarity and substructure queries), Reaction Group retrieval, Reaction Instance access, and Source Reference retrieval.

Methodology in Plain English

The authors start from a large pool of patent records aggregated by the Sciverse scientific data infrastructure, drawn from the USPTO and EPO and primarily covering organic synthesis patents published from 1976 through 2025. They apply successive automated stages: filtering by IPC classification for organic chemistry relevance, consolidating and deduplicating records within each source, deduplicating across sources, and removing records without usable full text. Retaining records conservatively when metadata did not support a reliable duplicate determination, this funnel leaves 1,580,939 patent documents entering extraction.

A fully automated information-extraction and normalization pipeline then parses patent text, images, and reaction schemes. At the capability level it identifies organic-synthesis experimental content and structures the participants, experimental conditions, yields, and experimental procedures of each specific experiment into a Reaction Instance, with canonical SMILES used for participants with usable structural information. In the data production described, DeepSeek-V3-0324 is used primarily to process patent experimental text and MinerU.Chem to parse chemical information in images and reaction schemes; DeepSeek-V3-0324 is deployed on Huawei Ascend 910C AI processors with DeepLink providing software–hardware adaptation and inference-runtime support. Each instance is linked back to its source patent and source-text location, and an automated qualification step using RXNMapper atom mapping and atom-conservation and other rules labels records as qualified or not passed. Both categories are retained.

For quality assessment, the authors treat the 14,808,205 qualified instances as the target population and randomly sample 1,300 records. Reviewers with professional backgrounds in organic chemistry independently compare each structured field with the corresponding content in the source patent for Yield, Reactant, Reagent, Catalyst, and Solvent; disagreements are resolved through re-evaluation and adjudication. For the Pistachio comparison, Pistachio is treated as an external comparator rather than ground truth: the authors analyze a matched sample of 100 US patents for deduplication, inspect a shared reaction from patent US10975080 for representation granularity, and identify 660 reaction pairs from 58 patents whose records belong to the same patent and have exactly matching normalized source paragraphs for field-level agreement scoring.

Why This Matters

The paper argues that existing open reaction datasets have substantial limitations in data scale, patent-literature coverage, instance-level experimental information, provenance localization, and consistency of data organization, and therefore cannot provide a complete, reliable, structurally consistent foundation for fine-grained reaction retrieval, experimental-condition comparison, source verification, and AI4Chem applications. Professional databases offer richer curated information and retrieval capabilities but typically require paid subscriptions or commercial licenses, with machine access, batch use, and system integration potentially subject to restrictions. DianShi-RxnDB is positioned against both gaps: comparable-scale fine-grained data with paragraph-level provenance, available through free non-commercial Web and MCP access.

Real-world applications implied by the paper:

  • Reaction-precedent retrieval and comparison of reported experimental procedures and conditions across patents.
  • Supplying data for reaction-outcome prediction, reaction-condition prediction, and data-driven retrosynthetic planning.
  • Source-patent verification: navigating from a structured record back to the relevant source paragraph to confirm procedures, roles, conditions, and yields.
  • Agent-driven workflows: multi-step retrieval over substances, reactions, and source documents composed through MCP tools, with object identifiers and source information that let users relocate records in the Web workbench.

Industry relevance: the platform targets pharmaceutical, materials, agrochemical, and fine chemical synthesis contexts, where organic synthesis is described as central to discovery and development. Its free non-commercial access model contrasts with the paid Web access listed for Pistachio, Reaxys, and CAS Reactions, and its dual researcher/agent interfaces address the emerging requirement that data platforms expose structured, composable tools rather than only downloadable files or human-facing search.

Future Directions

  • Reducing the analyzed error categories: the authors state that the observed errors identify directions for further data-quality improvement, including the treatment of workup information, participant-role assignment, cross-paragraph extraction, and multi-step experimental descriptions. Reagent classification (84.62% in the manual evaluation) is noted as especially prone to error because the reagent boundary depends strongly on reaction context.
  • A subsequent technical report: this report presents only a capability-level overview of the extraction pipeline, whose specific technical details are to be described separately.
  • Extending evaluation beyond the matched samples: the 1.368-times deduplication ratio is described as a property of the 100-patent matched sample and is not extrapolated to the complete corpora, and the manual evaluation covers qualified instances only, with results explicitly not applying to instances that did not pass automated qualification.
  • Boundaries, risks, and availability: Section 6 is described as discussing data and use boundaries, risks associated with agent use, and availability, and Section 7 as concluding the report and looking ahead to future work.

Target Audience

Organic chemistry researchers who need reaction-precedent retrieval, experimental-condition comparison, and source-patent verification; AI4Chem and cheminformatics researchers who need large-scale, fine-grained, provenance-linked reaction data for outcome prediction, condition prediction, and retrosynthetic planning; and AI agent developers building multi-step chemical information workflows over structured retrieval tools. Reviewers and practitioners evaluating reaction data resources against commercial alternatives such as Pistachio, Reaxys, and CAS Reactions will also find the comparison methodology and access-model discussion relevant.

Authors’ abstract

High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .

Read the original paper