Research
DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
Overview Research area: AI for education, specifically automatic evaluation of mathematical diagrams produced as LaTeX TikZ code by large language models. Technical level: Intermediate. The paper assu

- arXiv
- 2511.08283
- Published
- 2025-11-11
- Authors
- Vishal Kumar, Shubhra Mishra, Rebecca Hao, Rizwaan Malik, David Broman, Dorottya Demszky
AI summary
Overview
Research area: AI for education, specifically automatic evaluation of mathematical diagrams produced as LaTeX TikZ code by large language models.
Technical level: Intermediate. The paper assumes familiarity with LLM benchmarking, Cohen's kappa, and the general idea of compiler-style intermediate representations, but the core idea is explained clearly.
Scope: The paper proposes DiagramIR, an automatic pipeline that converts TikZ diagram code into a structured intermediate representation (IR) and applies rule-based mathematical and spatial checks, and shows this beats LLM-as-a-Judge on agreement with human raters.
What This Paper Is About
LLMs are increasingly generating educational math figures as compilable code, but there is no scalable way to check whether those figures are mathematically and visually sound. The authors build an evaluation pipeline that translates TikZ code into a schema-constrained intermediate representation and then runs deterministic checks on it, rather than asking a model to judge the diagram directly. The goal is an evaluation method accurate enough for live educational tools and cheap enough to run at scale.
Key Contributions
- A back-translation pipeline that uses an LLM to convert TikZ code into an intermediate representation, on which programmatic mathematical and spatial checks are run, outperforming LLM-as-a-Judge on agreement with human evaluation across all four tested models.
- A 398-item evaluation dataset of real-world mathematical diagram generations from teachers, drawn from 6,000 random conversations from Coteach (an AI assistant for mathematics educators using the Illustrative Mathematics K–12 Math v.360 curriculum), comprising 208 2D diagrams (triangles, circles, rectangles) and 190 3D shape diagrams (prisms and cubes).
- A six-criterion rubric covering mathematical correctness and spatial correctness, used for human annotation and mirrored in the automated pipeline.
- A demonstration that decoupling perception (TikZ to IR) from verification (rule-based checks) lets a small model such as GPT-4.1-Mini match a frontier model such as GPT-5 at substantially lower cost.
Main Findings
- Back-translation beats LLM-as-a-Judge on human agreement: Across four models, back-translation achieved Cohen's kappa of 0.48–0.56, while LLM-as-a-Judge in its strongest setting (code plus image) reached 0.39–0.47.
- Per-model kappa values (back-translation vs. LLM-as-a-Judge with code and image, Table 2): GPT-4.1: 0.562 vs. 0.399; GPT-5: 0.555 vs. 0.498; GPT-4.1 Mini: 0.483 vs. 0.388; GPT-5 Mini: 0.527 vs. 0.465.
- Cost and latency (Table 2, code plus image judge): Back-translation cost $6.75 (GPT-4.1), $10.29 (GPT-5), $0.47 (GPT-4.1 Mini), $2.12 (GPT-5 Mini); the judge cost $3.61, $4.85, $0.82, and $0.86 respectively. Average evaluation time under back-translation was 25.78s, 36.46s, 12.16s, and 42.14s.
- Small model parity at lower cost: GPT-4.1-Mini under back-translation performed comparably to the best LLM-as-a-Judge (GPT-5) at a reported 10.3x lower cost ($0.47 vs. $4.83 as stated in the text; Table 2 lists $4.85 for GPT-5 under the code-plus-image judge).
- Where each method wins (Table 3, image plus code judge): Back-translation was stronger for "Diagram fully in frame" (e.g., 0.604 for GPT-5 vs. 0.390), "Elements scaled to be readable" (e.g., 0.334 for GPT-5 vs. 0.043), and "No problematic overlap" (e.g., 0.608 for GPT-5 vs. 0.315). LLM-as-a-Judge was markedly stronger on angle labels (0.829 for GPT-5 Mini vs. 0.652 under back-translation) and lengths/areas (0.673 for GPT-5 vs. 0.429 for GPT-5 under back-translation).
- Confusion analysis (Tables 8–11, N = 386 diagrams, N = 2348 applicable slots): For GPT-5, totals were 152 true positives (6.5%), 1888 true negatives (80.4%), 185 false positives (7.9%), and 123 false negatives (5.2%). Comparable totals were reported for GPT-4.1, GPT-5-mini, and GPT-4.1-mini.
- Dataset error distribution (Table 12, human evaluation): For example, 345 (89.4%) diagrams were evaluated as fully in canvas versus 41 (10.6%) not; 297 (77.0%) had no problematic overlap versus 89 (23.1%); 369 (95.6%) had readable element sizes versus 17 (4.4%).
- Other judge settings (Appendix Tables 4 and 5): With code only, judge kappa was 0.395 (GPT-4.1), 0.427 (GPT-5), 0.388 (GPT-4.1 Mini), 0.406 (GPT-5 Mini); with image only, 0.365, 0.442, 0.366, and 0.442.
Methodology in Plain English
The pipeline works in two stages. First, an LLM reads the LaTeX TikZ code for a diagram and translates it into an intermediate representation: a structured JSON-like schema (the TikzIR class, containing lists of shapes, line segments, nodes, circles, rectangle primitives, arcs, clips, and picture options) that records geometric entities and their coordinates in a standardized form. This step is called back-translation, borrowing the idea from neural machine translation of projecting from a messy target language into a cleaner source language. Second, a set of rule-based checks operates on that IR. The six checks test whether labeled lengths or areas match drawn proportions, whether labeled angles match drawn angles, whether the diagram is fully in frame (with a default 2pt buffer), whether elements are large enough to read (relative threshold default 0.02 of the diagram's smaller bounding-box dimension), whether labels are attached to the correct elements, and whether elements overlap problematically (text-text overlap above 0.05 of the smaller box area, or a boundary crossing more than 0.4 of a label's perimeter).
The authors compare this against an LLM-as-a-Judge baseline in which the model is asked to apply the same rubric, given either the diagram image, the TikZ code, or both. All judge runs used temperature 0 and top-p 1, and GPT-5 was run in "low reasoning" mode to control latency. Human raters scored the same diagrams against the rubric, and Cohen's kappa measured agreement between each automated method and the human scores. Diagrams for human calibration and pipeline development (12 of them) were held out, leaving 386 for the test set.
Why This Matters
Impact on research: The paper argues for task-specific, symbolic-plus-lightweight-inference evaluation pipelines rather than relying on ever-larger frontier models for judging visual mathematical content. It also provides a grounded real-world dataset of teacher-requested diagrams, which the authors note is intended to support continued development of automatic evaluation for mathematical diagrams.
Real-world applications:
- Live math tutoring or co-teaching assistants that need to check generated figures before showing them to a teacher or student.
- Curriculum-aligned tools such as those built on the Illustrative Mathematics K–12 Math v.360 curriculum, where diagram requests come from real classroom practice.
- Low-resource or cost-constrained educational deployments, where the pipeline's lower cost for small models matters for feasibility.
- Auditing and quality control of diagram-generation systems, since the rule-based checks give explicit reasons why a diagram passes or fails.
Industry relevance: The cost comparison is the central commercial argument: if a small model plus deterministic checks can match a frontier model judge, educational technology providers can deploy diagram validation at a fraction of the inference budget. The paper explicitly frames this as important for equitable, scalable education technology.
Future Directions
- Extending the rubric beyond mathematical and spatial correctness to pedagogical usefulness, which the authors identify as a critical but more subjective dimension they did not cover.
- Expanding the intermediate representation schema and checks to handle more complex diagrams, such as multi-step constructions and coordinate plots, which the current IR does not fully capture.
- Fine-tuning a small model specifically for TikZ-to-IR translation, which the authors suggest could reduce cost further and mitigate the stochastic errors introduced by using LLMs for the parsing step.
- Validating the pipeline on other domains and diagram types, such as physics diagrams and freehand sketches, and integrating the method directly into diagram-generation tools.
Target Audience
Researchers and practitioners working on AI in education, LLM evaluation methodology, and code-generation-to-visualization systems. It is also relevant to engineers building math tutoring or co-teaching products who need a cheap, auditable way to validate generated diagrams, and to anyone interested in hybrid symbolic-plus-neural evaluation pipelines.
Authors’ abstract
Large Language Models (LLMs) are increasingly being adopted as tools for learning; however, most tools remain text-only, limiting their usefulness for domains where visualizations are essential, such as mathematics. Recent work shows that LLMs are capable of generating code that compiles to educational figures, but a major bottleneck remains: scalable evaluation of these diagrams. We address this by proposing DiagramIR: an automatic and scalable evaluation pipeline for geometric figures. Our method relies on intermediate representations (IRs) of LaTeX TikZ code. We compare our pipeline to other evaluation baselines such as LLM-as-a-Judge, showing that our approach has higher agreement with human raters. This evaluation approach also enables smaller models like GPT-4.1-Mini to perform comparably to larger models such as GPT-5 at a 10x lower inference cost, which is important for deploying accessible and scalable education technologies.