Research
SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
Overview Research area: Computer Vision / Multimodal Large Language Models (scientific image verification and reward modeling). Technical level: Advanced (assumes familiarity with supervised fine-tuni

- arXiv
- 2609.33399
- Published
- 2026-09-27
- Authors
- Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin, Xi Yu, Xusen Hei, DingBa Fu, Jiayuan Xie, Yi Cai
AI summary
Overview
Research area: Computer Vision / Multimodal Large Language Models (scientific image verification and reward modeling). Technical level: Advanced (assumes familiarity with supervised fine-tuning, reinforcement learning with GRPO, and multimodal benchmarks). Scope: The paper introduces SciGen-Verify, a benchmark for explainable verification of scientific image generation, and SciGen-Verifier-8B, a reasoning-driven multimodal verifier trained with a two-stage reinforcement learning curriculum.
What This Paper Is About
Generative models can now produce scientific diagrams such as circuits, geometric constructions, and function plots, but nothing reliable checks whether those diagrams are actually scientifically correct. Existing verifiers are built for natural images and reduce their output to a single score, so they neither cover scientific content nor explain what went wrong or how to fix it. This paper builds both a benchmark for this verification task and a compact model that outputs a binary judgement, a natural-language explanation, and an actionable editing instruction.
Key Contributions
- SciGen-Verify benchmark. The authors construct what they describe as the first benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge. It contains 1,350 data samples (406 instruction following, 287 multidisciplinary reasoning, 657 world knowledge) and uses a three-tier hierarchical protocol over binary judgement, supporting explanation, and corrective editing instruction.
- SciGen-Verifier-8B model. An 8B multimodal verifier built on Qwen3-VL-8B-Instruct, trained via cold-start supervised fine-tuning on 48,311 curated training instances, followed by a curriculum-based two-stage reinforcement learning pipeline using GRPO.
- Curriculum-based RL recipe. Stage 1 uses a rubric-guided process reward that checks whether the reasoning trace covers instance-specific scientific checkpoints; Stage 2 adds an outcome reward that aligns the judgement, explanation, and editing instruction with ground-truth annotation.
- Empirical and deployment results. The model reaches competitive performance against much larger proprietary baselines on SciGen-Verify and is demonstrated as an online critic for iterative image rectification through test-time scaling.
Main Findings
- Frontier models struggle with explainable verification. Even the strongest baseline, Qwen3.8-Max, reaches 86.03% overall Acc_a. GPT-6-Astra falls from 75.71% overall Acc_a to 53.59% Acc_i, and Gemini-3.5-Flash falls from 81.38% to 56.54%, indicating weak fine-grained mistake localization and correction.
- Generic verifiers do not transfer. OmniVerifier-7B, trained for evaluating generic images, performs poorly on SciGen-Verify (58.73% overall Acc_a, 45.99% Acc_e, 9.39% Acc_i), which the authors take as evidence that scientific verification demands disciplinary knowledge beyond text-image alignment.
- SciGen-Verifier-8B is competitive with far larger models. It achieves the highest overall Acc_a of 86.28%, above Qwen3.8-Max (86.03%), Gemini-3.5-Flash (+4.90%) and GPT-6-Astra (+10.57%). Its overall Acc_e is 81.79% and overall Acc_i is 57.42%.
- World knowledge is the strongest domain. SciGen-Verifier-8B reaches 94.21% Acc_a, 91.53% Acc_e, and 65.48% Acc_i there, exceeding Qwen3.8-Max by +6.45%, +4.41%, and +5.67% respectively. Its instruction-following scores are 75.62 / 70.38 / 48.93 and its reasoning scores are 89.02 / 83.47 / 57.84.
- Correcting errors is harder than detecting them. Acc_i drops sharply for all models. Gemini-3.5-Flash and GPT-6-Astra cap at 56.54% and 53.59% overall Acc_i, Qwen3.8-Max reaches 64.75%, and SciGen-Verifier-8B obtains 57.42%. The paper notes that Acc_i is normalized over negative samples only and is a conditional metric, not directly comparable to Acc_a and Acc_e.
- Chain-of-thought is a prerequisite for RL to work. Without CoT, SFT gives 81.0% Acc_a but collapses on feedback (65.9% Acc_e, 36.4% Acc_i); adding RL to that paradigm gives negligible gains. With CoT, SFT lifts Acc_e and Acc_i to 75.3% and 46.7% with a marginal Acc_a concession, and CoT plus RL reaches 86.3% Acc_a and 57.4% Acc_i.
- The two-stage curriculum outperforms single-stage training. SFT alone yields 80.09 / 75.32 / 46.72 overall; SFT plus outcome reward only yields 82.68 / 78.46 / 52.78 with reasoning Acc_a capped at 84.9% and Acc_i at 51.2%; SFT plus stage-1 RL reaches 83.62 / 77.16 / 50.47 with reasoning Acc_a at 86.1%; the full curriculum reaches 86.28 / 81.79 / 57.42.
- Judges are validated. Two model-based judges were checked against three human experts on 300 sampled instances, achieving agreement κ ≥ 0.76 against a human–human ceiling of κ = 0.83.
- Qualitative gains. In the case study, Seed-2.0-Pro produces a false positive by overlooking requested dotted lines, while SciGen-Verifier localizes the violation and produces an editing instruction. As an online critic in iterative test-time scaling, it diagnoses geometric mislabeling and mathematical inconsistencies to guide self-correction.
Methodology in Plain English
The authors first assembled data from existing scientific and educational sources: V-Interaction and MathCanvas for instruction following; CoSyn, DaTikZ-V3, EMMA, and crawled exam papers for multidisciplinary reasoning; and mined Wikipedia content plus AI2D for world knowledge. Because these sources differ in format, models such as Claude-Sonnet-5, Claude-Opus-4.8, and Gemini-3-Flash were used to generate missing natural-language instructions or captions.
To create wrong answers, they built negative samples in two ways. Visual perturbation subtly alters the rendered image, either by editing the rendering code or by having Gemini-3-Pro write an editing instruction that Qwen-Image-Edit applies to the original image. Textual perturbation instead changes the problem statement—reversing structural logic such as turning a parallel circuit into a series one, changing numerical values, or substituting chemical functional groups—while keeping the correct image. A filtering pass with Gemini-3-Flash removes blurry, illegible, or clearly botched candidates.
Annotation was initialized by Gemini-3.1-Pro and then reviewed by 10 domain experts, two graduate students for each of five disciplines (Mathematics, Physics, Chemistry, Biology, Geography). Paired positive and negative samples were split into separate groups, and a sample was kept only if both expert reviewers approved it.
The model itself starts as Qwen3-VL-8B-Instruct and is trained in stages. First, full-parameter supervised fine-tuning teaches a fixed XML-style output with reasoning trace, judgement, explanation, and edit fields, using rejection sampling to keep only correct judgements with validated explanations. Second, reinforcement learning with GRPO runs in two phases: stage 1 rewards how many instance-specific scientific rubric checkpoints the reasoning trace covers, with the weight shifting from the rubric reward toward the judgement reward as accuracy improves; stage 2 adds outcome rewards for semantically equivalent explanations and editing instructions, judged by GLM-4.7 and Doubao-Seed-2.0-Lite respectively.
Why This Matters
Verification is the step that determines whether AI-generated scientific diagrams can be trusted in education and research, and this paper argues the field has been measuring the wrong thing by reducing verification to a scalar score on natural images. By separating detection, explanation, and correction into cascading metrics, it exposes how much harder correction is than detection.
Real-world applications:
- Automated grading or assistant grading of diagram-based homework in mathematics, physics, chemistry, biology, and geography.
- Quality control for AI-generated scientific figures before publication or classroom use.
- Intelligent tutoring systems that explain why a student's or a model's diagram is wrong and how to fix it.
- Iterative image rectification pipelines where a verifier acts as an online critic inside a generative loop.
Industry relevance: the work targets education technology and generative AI quality assurance, showing that a small 8B model can approach or exceed much larger proprietary systems on a specialized verification task.
Future Directions
- Improving Acc_i, which remains the bottleneck: even the best baseline tops out at 64.75% and SciGen-Verifier-8B reaches 57.42%.
- Extending the curriculum-based RL recipe to other backbone models and sizes beyond the single 8B Qwen3-VL initialization used here.
- Broadening coverage beyond the five target disciplines and the three current benchmark domains, and testing generalization to unseen scientific subjects.
- Formalizing the online-critic loop for iterative test-time scaling, which is demonstrated qualitatively but whose broader scaling behavior is not reported.
Target Audience
Researchers and practitioners working on multimodal LLMs, reward models, and RL fine-tuning for vision-language tasks; benchmark designers in scientific and educational AI; and applied teams in education technology or generative image quality control who need verifiable, explainable feedback rather than a single score.
Authors’ abstract
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.