Research
ChemPro: A Progressive Chemistry Benchmark for Large Language Models
Overview Research area: Natural language processing — evaluation and benchmarking of large language models, specifically on scientific reasoning in chemistry. Technical level: Intermediate. The benchm

- arXiv
- 2602.03108
- Published
- 2026-02-03
- Authors
- Aaditya Baranwal, Shruti Vyas
AI summary
Overview
- Research area: Natural language processing — evaluation and benchmarking of large language models, specifically on scientific reasoning in chemistry.
- Technical level: Intermediate. The benchmark design and grading rules are explained in accessible terms, though some familiarity with LLM evaluation (pass@1, exact match, tolerance scoring) helps.
- Scope: ChemPro is a 4,100-question, curriculum-aligned chemistry benchmark spanning four difficulty tiers and four chemistry subfields, used to measure how 45+7 state-of-the-art LLMs degrade as question formulation grows more complex.
What This Paper Is About
Existing chemistry benchmarks tend to target undergraduate-to-expert research tasks, or they label difficulty subjectively, leaving foundational high-school chemistry under-evaluated. The authors build ChemPro, a progressive benchmark whose difficulty tiers are tied to real educational sources (NCERT textbooks, JEE Mains exams, web quizzes) so that "harder" means something verifiable rather than annotator opinion. The goal is to diagnose where and why LLMs fail as chemistry problems require longer chains of reasoning, not just whether they know chemistry facts.
Key Contributions
-
A progressive, source-aligned chemistry benchmark. ChemPro contains 4,100 natural-language question-answer pairs across four sections — Easy (𝒞𝒫_E), Medium (𝒞𝒫_M), Challenging (𝒞𝒫_C), and Difficult (𝒞𝒫_D) — with a curriculum-verified difficulty ordering (𝒞𝒫_E ≺ 𝒞𝒫_M ≺ 𝒞𝒫_C ≺ 𝒞𝒫_D). Each section is drawn from a distinct provenance: web sources and quizlets (𝒞𝒫_E), NCERT grades 9–10 (𝒞𝒫_M), NCERT grades 11–12 (𝒞𝒫_C), and JEE Mains 2020–2024 (𝒞𝒫_D).
-
Dual assessment modes. Every section partitions into multiple-choice questions and numerical problems, so the benchmark tests both conceptual understanding (MCQs, tolerance-based scoring) and computational precision (numericals, exact match).
-
Three-fold verification and leakage auditing. Each question passes source verification, expert review, and AI-assisted validation. Cross-benchmark deduplication using n-gram similarity found overlap of only 6 questions (0.15% of the dataset), and a four-method GPT-4o leakage analysis (prefix completion, paraphrasing, content modification, reverse engineering) identified roughly 8% potential exposure.
-
A large-scale evaluation. 45+7 LLMs are evaluated — 40 open-source models across five parameter scales (7B, 10B, 14B, 32B, 70B) plus 5 chemistry-corpus-pretrained models and 2 recent general-purpose releases, alongside proprietary systems (GPT-3.5-Turbo, GPT-4o, o1-mini, o3-mini, o1) and the ChemCrow agentic framework.
Main Findings
-
Consistent degradation across every model tested: All 45 evaluated models show declining accuracy as section difficulty increases, with an average of approximately 21 percentage points dropped from elementary to competitive formulations.
-
A 13-point gap between curriculum-equivalent sections: Because JEE Mains officially follows the NCERT syllabus, 𝒞𝒫_C and 𝒞𝒫_D are argued to cover identical conceptual scope. The average 13-point accuracy drop from 𝒞𝒫_C to 𝒞𝒫_D across all models is presented as evidence that difficulty comes from formulation complexity (Multi-step reasoning, unit conversions, cross-condition integration) rather than new concepts.
-
Scaling helps on easy questions only: Larger models consistently achieve ≥90% accuracy on 𝒞𝒫_E MCQs, and performance converges in the intermediate 𝒞𝒫_M and 𝒞𝒫_C sections. On competitive questions, proprietary reasoning models (o1, o3-mini) reach 74–76% accuracy, 70B+ variants reach 68–71%, and 7–10B models reach 53–67%.
-
Best-in-class results by size category (MCQ accuracy, 𝒞𝒫_E / 𝒞𝒫_M / 𝒞𝒫_C / 𝒞𝒫_D): Falcon3 7B Instruct 0.92 / 0.94 / 0.70 / 0.58 (7B); Falcon3 10B Instruct 0.96 / 0.97 / 0.77 / 0.62 (10B); Luminis-PHI-4 0.97 / 0.98 / 0.85 / 0.74 (14B); Rombos-LLM-V2.5-Qwen-32B 0.96 / 0.99 / 0.81 / 0.69 (32B); Rombos-LLM-V2.5-Qwen-72B 0.97 / 0.99 / 0.83 / 0.71 (70B); OpenAI o1 0.97 / 0.99 / 0.85 / 0.76 (proprietary).
-
Agentic scaffolding does not close the gap: ChemCrow scores 97 / 98 / 84 / 68 across the four MCQ sections versus GPT-4o's 95 / 98 / 84 / 71 — comparable or slightly worse at the hardest tier, suggesting tools and prompting do not resolve the underlying reasoning bottleneck.
-
Humans are far ahead, with a caveat: Historical JEE Mains and NCERT board data from 2020–2024 show top 100 students typically score 97–100% on chemistry sections. The authors explicitly frame this as an indirect proxy: the cited cohort never took ChemPro, and the sampling may not match any single historical paper.
-
Subfield-specific bottlenecks (MCQ accuracy, mean ± spread): Biochemistry is strongest (98.7 ± 1.2 at 𝒞𝒫_E, falling to 71.2 ± 8.1 at 𝒞𝒫_D) but is held back by computational demands; Organic Chemistry collapses at 𝒞𝒫_C (58.5 ± 6.1) and 𝒞𝒫_D (50.1 ± 7.4), tied to spatial reasoning and multi-step synthesis; Physical Chemistry handles mathematical formulation but fails on precise calculation and unit conversion; Inorganic Chemistry is the most variable (90.5 ± 4.3 at 𝒞𝒫_E down to 59.1 ± 9.2 at 𝒞𝒫_D).
-
Numerical answers are harder than conceptual ones: On tolerance-based scoring, 𝒞𝒫_D accuracy falls to 60.0 ± 12.3 (Bio), 47.8 ± 11.7 (Inorganic), 45.5 ± 9.8 (Organic), and 52.3 ± 10.2 (Physical), with exact-match scoring reported separately and plotted in the figures.
-
Four recurring failure modes: Breakdown on problems requiring more than three sequential logical steps; arithmetic errors in stoichiometry and equilibrium despite correct setup; failure to integrate information across a problem statement, especially in organic reaction mechanisms; and systematic unit-conversion/dimensional-analysis mistakes.
-
Dataset skew toward hard questions: 𝒞𝒫_D contains 2,315 items (56.4%), 𝒞𝒫_C 665 (16.2%), 𝒞𝒫_M 335 (8.2%), and 𝒞𝒫_E 795 (19.3%). The paper states a total of 4,100 questions; as listed, these section counts sum to 4,110.
Methodology in Plain English
The authors did not write new chemistry questions from scratch. They collected existing questions from trusted educational sources — online quizzes for the easiest tier, NCERT textbooks for grades 9–10 and 11–12 for the middle tiers, and JEE Mains competitive exam papers from 2020 to 2024 for the hardest tier. Because those sources already encode decades of curriculum consensus about what is easy and what is hard, difficulty labels come from provenance rather than individual annotators.
Every question then passes three checks: where it came from, review by a human expert, and automated consistency checks for formatting, duplication, and leakage. Since chemistry relies heavily on diagrams, structures, and mechanisms, the authors converted visuals into text using standardized LaTeX equations, IUPAC nomenclature, numbered reaction steps, and structured diagram descriptions.
For evaluation, MCQs are scored by whether the model picks the correct letter. Numerical problems are scored two ways: exact match, which requires the precise value, and tolerance-based scoring, which accepts an answer within a threshold of θ = 0.1 (10%) of the average of the predicted and true values. This split separates conceptual understanding from computational precision. All models run pass@1, averaged over 5 runs, with a token budget of 8,000 for standard LLMs and 10,000 for reasoning models, temperature 0.3 for LLMs and 1 for reasoning models, and top-p 0.9, on two Nvidia A100 80GB GPUs. Questions ask for a solution under a "SOLUTION:" heading and the final answer under a "FINAL ANSWER:" heading, and numericals were parsed with a regex-based extractor that handles scientific notation.
Why This Matters
Impact on research: The paper reframes LLM difficulty as a property of articulation — the composition of required operations such as chained steps, conversions, and cross-condition integration — rather than vocabulary or topic coverage. Because 𝒞𝒫_C and 𝒞𝒫_D share a syllabus, the 13-point gap isolates formulation complexity as a measurable variable, giving researchers a diagnostic axis that parameter scaling and tool use do not address.
Real-world applications:
- Evaluating AI tutors and homework-help systems for high-school and early-undergraduate chemistry.
- Quality assurance for educational content platforms that use LLMs to generate or grade chemistry problems.
- Screening models before deploying them in scientific research assistants, where silent arithmetic and unit errors are costly.
- Designing curriculum-aligned adaptive learning systems that know which reasoning skills a student (or a model) has not yet mastered.
Industry relevance: Chemistry is described as the "central science," and the authors note that chemistry has received comparatively less benchmark attention than mathematics or programming — domains with formal languages that make verification easier. Companies building domain-specific scientific AI, education technology platforms, and agentic lab assistants all need to know whether their system degrades gracefully or silently on multi-step chemistry problems, and ChemPro provides a tiered way to find out. The finding that ChemCrow performs no better than its base GPT-4o model is directly relevant to anyone betting on tool-augmented agents as a fix.
Future Directions
-
Reasoning robustness as a training target. The paper concludes that current architectural paradigms cannot overcome complex scientific reasoning through scaling or agentic frameworks alone, implicitly calling for methods that specifically target multi-step composition, unit handling, and cross-condition integration.
-
Closing the subfield gaps. Organic chemistry spatial reasoning and multi-step synthesis, physical chemistry numerical precision, and the high variance in inorganic chemistry are identified as distinct bottlenecks deserving targeted investigation rather than a single general fix.
-
Stronger leakage and human baselines. The authors flag their human comparison as an indirect proxy since the cohort never took ChemPro end-to-end. A controlled human-vs-model head-to-head on the same items, alongside continued leakage auditing beyond the ~8% potential exposure found, would sharpen the benchmark's claims.
-
Extending the progression. ChemPro deliberately stops at high-school boundaries (E–HS). Whether the same articulation-based difficulty accounting holds for undergraduate and graduate material — where existing benchmarks like ChemBench and GPQA operate — remains open.
Target Audience
Researchers working on LLM evaluation and scientific reasoning; developers building chemistry or STEM-focused models and agents; education technology teams deploying AI for tutoring, problem generation, or grading; and educators or curriculum designers interested in how model failure modes map onto the specific skills students are expected to master at each grade level.
Authors’ abstract
We introduce ChemPro, a progressive benchmark with 4100 natural language question-answer pairs in Chemistry, across 4 coherent sections of difficulty designed to assess the proficiency of Large Language Models (LLMs) in a broad spectrum of general chemistry topics. We include Multiple Choice Questions and Numerical Questions spread across fine-grained information recall, long-horizon reasoning, multi-concept questions, problem-solving with nuanced articulation, and straightforward questions in a balanced ratio, effectively covering Bio-Chemistry, Inorganic-Chemistry, Organic-Chemistry and Physical-Chemistry. ChemPro is carefully designed analogous to a student's academic evaluation for basic to high-school chemistry. A gradual increase in the question difficulty rigorously tests the ability of LLMs to progress from solving basic problems to solving more sophisticated challenges. We evaluate 45+7 state-of-the-art LLMs, spanning both open-source and proprietary variants, and our analysis reveals that while LLMs perform well on basic chemistry questions, their accuracy declines with different types and levels of complexity. These findings highlight the critical limitations of LLMs in general scientific reasoning and understanding and point towards understudied dimensions of difficulty, emphasizing the need for more robust methodologies to improve LLMs.