Research
EDUMATH: Generating Standards-aligned Educational Math Word Problems
Overview Research area: Natural Language Processing — LLM-based educational content generation, math word problem generation, human-in-the-loop evaluation. Technical level: Intermediate. The paper ass
- arXiv
- 2510.06965
- Published
- 2025-10-08
- Authors
- Bryan R. Christ, Penelope Molitz, Beau LeBlond, Zachary Gottesman, Jonathan Kropko, Thomas Hartvigsen
AI summary
Overview
Research area: Natural Language Processing — LLM-based educational content generation, math word problem generation, human-in-the-loop evaluation.
Technical level: Intermediate. The paper assumes familiarity with LLMs, supervised finetuning, preference optimization, and text classification, but the evaluation design and educational framing are explained in plain terms.
Scope: The paper introduces EDUMATH, a pipeline of teacher-annotated data and finetuned open models for generating grade 3-5 math word problems that are aligned to specific US math standards and customized to student interests, evaluated against open and closed baselines and in a classroom study.
What This Paper Is About
Math word problems are a core K-12 teaching tool, and research suggests personalizing them to a student's interests and ability level helps learning. Teachers, however, rarely have time to write custom problems, so they fall back on generic question sets that are not tailored to individuals and are not easily searched by math standard. This paper asks whether LLMs can generate math word problems that simultaneously satisfy a specific math standard (for example, single-step whole-number multiplication), include a correct and readable solution, and are customized to what a student finds interesting — and whether students actually benefit.
Key Contributions
-
A benchmark gap analysis. The authors find that open models, especially small ones, struggle to generate standards-aligned educational math word problems relative to closed models, and report performance gaps between open and closed models and between smaller and larger open models.
-
The first teacher-annotated training dataset for this task. Real teachers annotated over 3,000 generated math word problems, which were filtered into the Standards-Targeted Educational Math dataset (STEM), described as the first training dataset for standards-aligned educational MWP generation.
-
Two state-of-the-art generators. The authors use their data to train EDUMATH 12B and EDUMATH 30B, both released along with the annotated data.
-
The first classroom study of customized LLM-generated math word problems with grade school students, testing both performance and preference.
Main Findings
-
Teacher annotation produced a high-quality but non-trivial labeling task. Across 1,372 US teachers on Prolific, the first two annotators agreed on solvability 90.1 ± 1.2% of the time, accuracy 76.6 ± 1.5%, educational appropriateness 77.5 ± 1.5%, standards alignment 75.3 ± 1.6%, and "meets all criteria" (MaC) 65.5 ± 1.7% of the time. The authors note the accuracy and MaC agreement rates are lower than those reported in Christ et al. (2024), which they attribute to accuracy here covering intermediate reasoning and solution readability, and to MaC spanning four criteria rather than three.
-
STEM is the only dataset in their comparison with teacher annotation, readable solutions, and full math standard annotation. STEM contains 2,577 deduplicated questions; the dataset also identifies 1,552 MWPs that met all criteria. Comparable datasets in the paper range from 1,000 (SVAMP) to 8,792 (GSM8K) questions. STEM's average question length is 53.9 tokens with readability 0.384 and BERTScore F1 73.6.
-
EDUMATH 30B has the highest MaC rate of any model evaluated, at 94.6 ± 0.9%, significantly better than the second-best open model at p < 0.01. Closed baselines GPT-4o and GPT-4.1 reach 92.8 ± 1.6% and GPT-4.5 reaches 92.0 ± 1.7%.
-
EDUMATH 12B closes most of the gap to much larger open models. It reaches 85.9 ± 1.2% MaC, above Gemma 3 27B IT (75.4 ± 1.3%) and close to Qwen 3 30B IT (87.3 ± 1.0%), while being far above its base model Gemma 3 12B IT (63.9 ± 1.5%).
-
A text classifier trained on the annotations lets an untrained 30B model beat closed baselines. Stacking the ModernBERT classifier on Qwen 3 30B IT yields EDUMATH 30B; the classifier itself reaches 79% accuracy and 0.861 AUC-ROC on a held-out test set.
-
EDUMATH outputs look more like human-written problems on automatic metrics. EDUMATH 12B has the lowest perplexity of all models at 9.5 (0.10), and its BERTScore relative to ASDIV (74.5) is closest to ASDIV's within-dataset BERTScore (73.6). EDUMATH 12B's average question length (54.9 tokens) is closest to STEM's and ASDIV's. Both EDUMATH models have the shortest solutions among open models (166.7 and 163.5 tokens).
-
Performance varies sharply by math topic. EDUMATH 12B is strongest on time problems (98.1) and weakest on fractions (78.6). EDUMATH 30B is strongest on measurement conversion (100.0) and weakest on multiplication/division (92.0), and is the only model scoring above 90% on every topic.
-
Students performed similarly on LLM-generated and human-written problems but preferred the generated ones. At the first school, 4th graders (n = 82) solved worksheets over four weeks. At the second, 3rd-5th graders (n = 12) worked with a math interventionist using individually customized problems; every student but one preferred the LLM-generated problem. Students most often said they liked it because of the topic.
-
The automated annotator matched human agreement levels. Expert annotators agreed with each other 76% of the time on MaC and agreed with the Gemma 3 27B IT annotator an average of 75% of the time; Cohen's kappa was 0.34 between annotators and 0.30 between annotators and the model.
Methodology in Plain English
The work proceeds in five phases.
Phase one: labeling human-written problems. No grade-school math dataset was annotated for math standards, so the authors took the 3rd-5th grade subset of ASDIV (1,027 problems) and labeled it for Virginia Standards of Learning (VA SOL) in four stages: an undergraduate education student labeled problems against all relevant standards, a team member with K-12 teaching experience corrected the labels, Llama 3.3 70B IT was prompted to check each label, and the teaching-experienced team member reviewed again. Two unsolvable problems were removed and several were rewritten. Gemma 3 27B IT reviewed generated step-by-step solutions and rewrote overly complex ones. The result was 1,025 problems with gold standard labels and chain-of-thought solutions.
Phase two: generating and teacher-annotating synthetic data. Because 1,025 problems was too few for training, they generated 3,012 problems with Llama 3.3 70B IT — roughly 90 per VA SOL combination missing from the labeled data and roughly 80 per combination present but with fewer than 100 samples. Teachers on Prolific labeled each problem twice on four criteria (a third time if the two disagreed), with final labels by majority vote. A problem "meets all criteria" only if all four pass. Gemma 3 27B IT independently labeled the same problems with the same directions; problems teachers passed but the model failed were flipped to failing to guard against human labeling errors, while problems teachers failed were kept as failing to preserve teacher expertise.
The four criteria are Solvability (the problem is answerable and has one correct answer), Accuracy (the chain-of-thought solution is correct and readable, so a correct answer with confusing reasoning counts as inaccurate), Educational Appropriateness (a teacher would be comfortable giving it to a 3rd-5th grader), and Standards Alignment (the problem genuinely exercises the specified standard). The first three are adapted from Christ et al. (2024); Standards Alignment is new.
Phase three: training. They finetune Gemma 3 12B IT on STEM with supervised finetuning for 5 epochs, then apply Kahneman-Tversky Optimization on 4,039 rows combining their annotations and the ASDIV subset. A ModernBERT classifier trained on 3,664 rows then filters the KTO model's outputs to yield EDUMATH 12B. Stacking the same classifier on Qwen 30B produces EDUMATH 30B.
Phase four: evaluating. They compare 1,000 problems from each EDUMATH model against 1,000 from each open model and 250 from each closed model (fewer due to API cost), using 8-shot prompting from STEM for every model so comparisons are fair.
Phase five: classroom study. Students solved one generated and one human-written problem per worksheet. Problems were screened by teachers beforehand. The first school used identical problems for all students in a grade over four weeks; the second used problems individually customized to each student based on a survey of their interests.
Why This Matters
Impact on research. The paper argues that prior educational math generation work either omits solutions, aligns only to broad topics like "multiplication" rather than the full language of a standard with its complexity constraints, or evaluates without teachers. It contributes a teacher-verified dataset, a reusable automated annotator, a trained text classifier, and a released set of models and annotations. It also reports a large gap between open and closed model performance on this task and shows that targeted data and lightweight post-filtering can substantially narrow it.
Real-world applications:
-
Teacher time savings. Generating standard-aligned problems on demand reduces the manual writing and curation load described in the paper's motivation.
-
Targeted practice. Fine-grained standard control (for example, single-step whole-number multiplication not exceeding 100) lets a teacher request exactly the skill a student needs.
-
Interest-based customization. Problems can be built around topics individual students like, which the classroom study links to student preference.
-
Intervention settings. The second school study involved students receiving math interventionist services, suggesting use in small-group or individualized instruction.
Industry relevance. The results are directly relevant to developers of educational technology and tutoring products, to teams building domain-specific finetuning and filtering pipelines, and to anyone weighing open versus closed models on a specialized generation task where a trained classifier can substitute for a larger model.
Future Directions
-
Extending beyond grades 3-5. The authors restricted the work to these grades because they are where math word problems are most used and because a narrow range allowed full coverage of standards under a fixed annotation budget; they note other grade levels are a compelling direction.
-
Multi-modal problems. The current problems are text-only, while many problems students encounter include images, tables, or figures.
-
Better prompting. A single standard prompt was used for all models to keep comparisons fair; the authors suggest examining prompt engineering to improve specific models.
-
Lower-cost evaluation. Human annotation cost is described as a broad limitation of MWP generator research, and the authors hope their annotations, classifier, and Gemma 3 27B IT annotator motivate automatic classification work.
-
Continued teacher involvement. The automated annotator may still miss nuances only experienced teachers identify, so future work should keep teachers in the evaluation loop.
-
Safety before deployment. The ethics statement notes EDUMATH may generate questions that are not educationally appropriate and that further research is needed before classroom deployment.
Target Audience
Researchers in NLP and educational data mining working on content generation and human-in-the-loop evaluation; AI and machine learning engineers building finetuning and output-filtering pipelines for specialized domains; education researchers and curriculum specialists interested in standards alignment; and edtech product teams evaluating whether open models plus targeted data can replace closed-model APIs for generating instructional content.
Authors’ abstract
Math word problems (MWPs) are critical K-12 educational tools, and customizing them to students' interests and ability levels can enhance learning. However, teachers struggle to find time to customize MWPs for students given large class sizes and increasing burnout. We propose that LLMs can support math education by generating MWPs customized to student interests and math education standards. We use a joint human expert-LLM judge approach to evaluate over 11,000 MWPs generated by open and closed LLMs and develop the first teacher-annotated dataset for standards-aligned educational MWP generation. We show the value of our data by using it to train a 12B open model that matches the performance of larger and more capable open models. We also use our teacher-annotated data to train a text classifier that enables a 30B open LLM to outperform existing closed baselines without any training. Next, we show our models' MWPs are more similar to human-written MWPs than those from existing models. We conclude by conducting the first study of customized LLM-generated MWPs with grade school students, finding they perform similarly on our models' MWPs relative to human-written MWPs but consistently prefer our customized MWPs.