Skip to content
AI.info

Research

Multi-Agent Collaborative Framework For Math Problem Generation

Multi-Agent Collaborative Framework For Math Problem Generation Overview Research area: Automatic Question Generation (AQG) for mathematics education, multi-agent LLM systems, and inference-time compu

Multi-Agent Collaborative Framework For Math Problem Generation
arXiv
2511.03958
Published
2025-11-06
Authors
Kia Karbasi, Kevin Hong, Mohammad Amin Samadi, Gregory Pottie

AI summary

Multi-Agent Collaborative Framework For Math Problem Generation

Overview

Research area: Automatic Question Generation (AQG) for mathematics education, multi-agent LLM systems, and inference-time computation (ITC) for Intelligent Tutoring Systems (ITS).

Technical level: Intermediate. The paper assumes familiarity with large language models, prompting strategies, and multi-agent workflows, but the framework itself is described at a conceptual level that is accessible to readers without deep expertise in agent architectures.

Scope: The paper proposes and evaluates two collaborative multi-agent workflows—Teacher-Critic Cycle (TCC) and Collective Consensus (CC)—for generating math questions and answers at controlled difficulty levels, evaluated with an automated GPT-4-based rubric.

What This Paper Is About

Intelligent Tutoring Systems need to produce practice problems on demand for any knowledge component (KC) at any requested difficulty, but existing transformer-based generators struggle to precisely control a question's complexity and cognitive demand. The authors introduce a collaborative multi-agent framework that uses inference-time computation—specifically, multiple agents that critique, revise, and converge on question-answer pairs—to better balance cognitive challenge against clarity and correctness. The goal is generated educational content that is more controllable and pedagogically useful than what a single model produces in one pass.

Key Contributions

  1. Two novel collaborative frameworks for AQG. The paper introduces Teacher-Critic Cycle (TCC), a two-agent iterative critique loop, and Collective Consensus (CC), a multi-agent discussion ending with a Consensus CEO that selects the final question-answer pair.

  2. A self-curation method guided by Bloom's taxonomy. A dedicated "Bloom Agent" scores each generated question from 1 to 5 on cognitive demand, mapping it to lower (Remembering, Understanding), middle (Applying, Analyzing), or upper (Evaluating, Creating) tiers. Candidates failing the expected cognitive challenge are aggressively discarded; this is compared against Random Curation (RC).

  3. An automated evaluation framework for generated math questions. Questions are scored on five criteria—clarity, relevance, importance, difficulty matching, and answerability—using prompt engineering based on state-of-the-art NLG evaluation frameworks (G-Eval and QGEval).

  4. Two research questions examined empirically: (RQ1) whether difficulty matching depends on the proposed difficulty level, and (RQ2) whether agentic workflows improve question generation compared to baseline models.

Main Findings

  • Curated agentic methods scored highest overall. TCC reached an average score of 4.64 and CC reached 4.65, compared with 4.40 for Baseline Teacher Zero-Shot (Baseline_ZS) and 4.42 for Baseline Teacher Few-Shot (Baseline_FS). The authors report that TCC and CC outperform both baselines and their non-curated counterparts across all evaluation metrics, but characterize the improvements as incremental rather than drastic.

  • Difficulty Matching improved most sharply. CC scored 4.96 and TCC 4.88 on difficulty matching, versus 4.41 (Baseline_ZS), 4.02 (Baseline_FS), 3.94 (CC_RC), and 4.11 (TCC-RC). Relevance was the other area of greatest gain, with CC at 4.99 and TCC at 4.92.

  • Non-curated agentic methods underperformed the baselines. CC_RC averaged 4.34 and TCC-RC averaged 4.45, both below the zero/few-shot baselines' 4.40 and 4.42. The authors conclude that agent-based generation alone is insufficient without structured selection or refinement, noting greater variability in agent responses.

  • Harder questions are harder to match. RQ1 is answered affirmatively: both difficulty matching and average score decrease as question difficulty increases (easy to medium to hard). Baseline models and non-curated agentic methods performed significantly worse on hard questions.

  • Zero-shot beat few-shot on difficulty matching. Unexpectedly, Baseline_ZS (4.41) scored higher on difficulty matching than Baseline_FS (4.02) and several agentic variants, suggesting few-shot examples may introduce biases or inconsistencies in aligning questions with intended difficulty.

  • More compute did not reliably help. Across Figures 4(a)-(c), increasing the number of rounds, the number of agents in CC, or the product of agents and rounds did not guarantee performance gains and could yield diminishing returns. The authors state that a more comprehensive parameter search is needed.

  • Prompting strategy had minimal effect. The three difficulty prompting strategies—empirical, prompting empirical, and prompting simple—produced nearly flat results (Table 2). For CC, difficulty matching ranged from 4.68 to 4.71 and average score from 4.60 to 4.64; for TCC, difficulty matching ranged from 4.61 to 4.67 and average score from 4.64 to 4.66. The authors conclude current few-shot learning strategies may be suboptimal for AQG.

  • Evaluation showed a ceiling effect. The GPT-4-based scores clustered near the top of the scale, similar to what the G-Eval authors reported, which can mask finer-grained differences between systems and limit the discriminative power of automated evaluation.

Methodology in Plain English

The system takes three inputs: a Knowledge Component name (per Common Core State Standards definitions), a set of example questions for that KC, and a required difficulty level (easy, medium, or hard). It outputs one generated question and its answer.

Data. Experiments use the Problem Bodies dataset, an extension of the ASSISTments dataset containing middle-school math questions with a "percent correct" attribute recording the percentage of students who answered correctly. This real-world performance metric was used to label questions as easy, medium, or hard (higher percent correct means easier). The paper does not report the number of questions in the dataset.

Agents. Four roles were defined: a Teacher that generates questions and answers for a specified KC and difficulty; a Generic Critic that gives high-level feedback on clarity, relevance, and difficulty alignment without adding new content; a Consensus CEO that reviews the conversation and selects the best pair; and a Versatile Agent that can generate a new pair, revise an existing one, or endorse a peer's pair with feedback.

Workflows. TCC pairs a teacher with iterative critic feedback over two to five interaction rounds. CC opens with one versatile agent generating a pair, then two to four versatile agents sequentially contribute by creating, revising, or explicitly agreeing; decoding parameters (sampling seed and temperature) are randomized per agent to encourage diverse perspectives, and after two to five rounds the Consensus CEO picks the final pair. Both workflows were tested with Auto Chain-of-Thought (step-by-step reasoning before the final answer) and explicit Solution Generation (output only the final answer) enabled and disabled.

Prompting conditions. Three variants were compared: Empirical (few-shot examples labeled easy/medium/hard using real student performance data), Prompting Empirical (only examples matching the requested difficulty are shown), and Prompting Simple (randomly selected examples from all difficulty tiers, so the model relies only on the stated difficulty).

Curation. The Bloom Agent scored each candidate 1-5 for cognitive demand, and candidates failing the expected challenge were discarded. This was benchmarked against Random Curation.

Evaluation. An automated GPT-4-based module scored each question from 1 to 5 on relevance, importance, clarity, difficulty matching, and answerability. Prompts and the evaluation module are available in the authors' repository (github.com/aminsmd/QA_GEN).

Why This Matters

Research impact. The paper connects two active lines of work—inference-time computation in LLMs and automatic question generation for education—and provides an evaluation protocol for a task where ground truth is subjective and human ratings are expensive. Its negative and mixed results (non-curated agents underperforming baselines, diminishing returns from more compute, flat prompting effects, and evaluation ceiling effects) are as informative as its positive ones, and they point to concrete gaps in current agentic and few-shot methods.

Real-world applications.

  • Intelligent Tutoring Systems that dynamically supply personalized practice problems for a given knowledge component at a requested difficulty, rather than relying on pre-authored problem banks.
  • Adaptive learning platforms that adjust a student's curriculum in real time based on an ongoing knowledge-tracing module.
  • Automated homework and worksheet generation for teachers, reducing the effort of writing new problem variations.
  • Content pipelines for educational publishers that need to scale question banks across many knowledge components and difficulty tiers.

Industry relevance. The paper was authored at the University of California, Los Angeles and the University of California, Irvine, and its target—scalable, difficulty-calibrated educational content—maps directly onto the needs of edtech companies, tutoring services, and assessment providers that currently depend on manual item writing or shallow variations of existing problems. The finding that structured refinement beats raw agent count is an important cost signal for anyone building agentic pipelines.

Future Directions

  1. Optimize inference-time computation. The authors note that merely adding rounds or agents does not guarantee gains and can produce diminishing returns, and call for a more comprehensive parameter search to identify where additional computation produces real benefit.

  2. Improve few-shot and in-context learning for AQG. Since the three prompting strategies produced little to no difference in performance, the authors suggest exploring alternative prompt engineering and adaptive few-shot learning techniques.

  3. Strengthen evaluation reliability. All findings rest on automated GPT-4 evaluations, which raises concerns about LLM self-evaluation bias and ceiling effects. The authors plan to collect human evaluation data and fine-tune the automated evaluation module to align more closely with human evaluators.

  4. Make harder questions better. Because difficulty matching degraded as difficulty increased, and baselines and non-curated agentic methods performed significantly worse on hard questions, structured prompting and iterative refinement specifically targeting high-difficulty calibration remain an open problem.

Target Audience

This paper is most useful to researchers and practitioners working on intelligent tutoring systems, educational data mining, and automatic question generation, as well as to machine learning researchers studying multi-agent collaboration and inference-time computation. It is also relevant to edtech product and engineering teams considering agentic LLM pipelines for content generation, and to educators or curriculum designers interested in the feasibility of automatically producing difficulty-calibrated practice problems. Readers looking for strong positive results should note the authors' own framing of the gains as incremental and their emphasis on the limitations of the automated evaluation.

Authors’ abstract

Automatic question generation (AQG) for mathematics education remains an elusive goal for Intelligent Tutoring Systems and educators. While pre-trained transformer-based language models have significantly advanced natural language generation, they often struggle to precisely control problem complexity and cognitive demands. In this paper, we introduce a collaborative multi-agent framework as a novel method of incorporating inference-time computation into AQG. This approach leverages multiple agents that iteratively refine generated question-answer pairs to better balance complexity and cognitive demand. We evaluate the generated questions on five meta-evaluation criteria: relevance, importance, clarity, difficulty matching, answerability, to assess the system's ability to control the required complexity and quality of the questions. Preliminary evaluations show that this collaborative multi-agent framework elevates the quality of generated educational content by fostering a more nuanced balance between cognitive challenge and clarity. These promising outcomes suggest that integrating collaborative multi-agent workflows can yield more controlled, pedagogically valuable content that can help advance automated educational content generation and adaptive learning environments.

Read the original paper