Skip to content
AI.info

Research

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Overview Research area: Large language model reasoning, specifically mathematical reasoning, benchmark design, and post-training (self-distillation). Technical level: Intermediate. The paper combines

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
arXiv
2610.02191
Published
2026-10-01
Authors
Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu

AI summary

Overview

Research area: Large language model reasoning, specifically mathematical reasoning, benchmark design, and post-training (self-distillation).

Technical level: Intermediate. The paper combines a new conceptual definition, a new benchmark, a diagnostic study of 12 models, and a training objective, so readers benefit from familiarity with LLM evaluation and fine-tuning, but the core argument is presented conceptually.

Scope: The paper argues that a model's ability to discover the load-bearing mathematical idea behind a problem is separable from its ability to execute a solution, and it introduces a benchmark (Prim) and a post-training method (Absorb) built around that distinction.

What This Paper Is About

Large language models are increasingly solving hard mathematics, but final-answer accuracy alone does not reveal whether a model actually grasped the structure that makes a problem solvable. The authors define a "Mathematical Primitive" — the concise, non-procedural insight that explains why a problem can be solved, such as an invariant, a reduction, or a reformulation — and build a benchmark to test whether models can find it, recover it from a solution, and use it. The goal is to diagnose where mathematical reasoning actually breaks down, and then to use that diagnosis to design a better post-training procedure.

Key Contributions

  1. A new conceptual object and benchmark. The authors introduce the Mathematical Primitive and Prim, a benchmark evaluating mathematical reasoning along four dimensions: Discovery, Generation, Digestion, and Execution. The final evaluation set contains 182 problems drawn from HLE-Verified.

  2. A systematic diagnosis of 12 models. The study covers three model lineages (OpenAI, Qwen, DeepSeek-R1 distill) and reports that answer accuracy masks distinct capability profiles, that correct primitives unlock substantial latent execution capacity, and that independent Discovery is the dominant bottleneck.

  3. A repairability analysis of failure modes. Using Supervised Fine-Tuning (SFT) and On-Policy Self-Distillation (OPSD), the authors show that "discovery-limited" failures are substantially more amenable to post-training than other failures, but that existing methods also introduce regressions.

  4. A new post-training paradigm, Absorb. Absorb treats the primitive as privileged information for the teacher and uses a bounded override mechanism to selectively transfer primitive-guided reasoning into the student, without requiring primitives at inference time.

Main Findings

  • Answer accuracy masks distinct capability profiles. Models with comparable Generation scores can differ sharply in Discovery. Qwen3.6-27B and gpt-5.4-mini achieve comparable Generation performance but differ by over 30 percentage points in Discovery. The OpenAI lineage is described as more "primitive-forward" (frequent correct Discovery with failed Generation), while Qwen shows more cases of failed Discovery with successful Generation, and DeepSeek-R1 distilled models concentrate in joint Discovery–Generation failures.

  • Correct primitives unlock latent execution capacity. Providing the gold primitive improves Execution by 17.58 to 29.67 percentage points across all 12 models. Under primitive guidance, Qwen models reach Execution performance competitive with the closed-source gpt-5.4 variants despite substantially weaker Discovery.

  • The gain is not just task decomposition or generic scaffolding. Conditioning a model on its own generated primitive gives little benefit (a marginal gain for gpt-5.4-nano, but a substantial degradation of Qwen3.6-27B, at -11.54). A step-by-step Teacher Plan improves over direct Generation, but a Teacher Primitive from the same teacher yields larger gains, and the Gold Primitive yields the most: 71.98%, 68.68%, and 78.57% for the three representative models.

  • Independent Discovery is the dominant bottleneck. Models perform much better on Digestion (recovering the primitive from a correct solution) than on Discovery (finding it from the problem alone). Qwen3.6-27B rises from 24.73% Discovery to 92.31% Digestion; R1-Distill-32B rises from 6.04% to 65.38%. A striking 83.6% of all Generation failures fall in the D− regime.

  • Discovery-limited failures are especially repairable. Across the 27B, 9B, and 4B models, the base models leave 341 failure instances in Prim, of which 294 have failed Discovery (147 D−E+ and 147 D−E−). SFT and OPSD repair 20.4% and 21.1% of D−E+ cases, roughly three times their rates of 6.8% and 8.2% for D−E− cases. This ordering holds at each model scale.

  • Existing post-training causes regressions. Net Generation gains stay modest because some initially correct cases become incorrect, consistent with post-training drift.

  • Absorb outperforms SFT and OPSD. SFT and OPSD reduce Generation accuracy by 6.04 and 4.95 percentage points at the 9B scale, and OPSD drops 3.30 points at 27B. Absorb improves Generation at all three scales, including a 5.49-point gain at 9B, and improves average benchmark performance by 2.42 to 4.48 points.

  • The benefit appears in reasoning, not in explicit primitive recovery. At 9B, Generation improves by 5.49 points while Discovery increases by only 1.10. At 27B, Generation improves by 1.10 points and HLE Math by 3.93 points, despite a 0.55-point decrease in Discovery.

Methodology in Plain English

The authors separate "structural understanding" from "procedural execution" and build a benchmark around that split.

  • Data curation: Starting from the HLE mathematics, text-only, free-form subset with exactMatch evaluation, they use the HLE-Verified Gold and Revision subsets and randomly sample 200 problems. GPT-5.4-High generates an initial primitive annotation conditioned on the problem, reference answer, and gold rationale. Three human experts with graduate-level mathematics training independently review each primitive and verify the problem, answer, and rationale. This excludes 18 unsuitable problems, yielding 182. For 17 of those 18, no valid gold primitive could be established. Expert review also identified residual errors in the official answers of 15 retained problems (13 mathematical errors, 2 corrupted answer strings), which were corrected; re-evaluating the strongest model on the corrected answers flipped 11 predictions from incorrect to correct, with no reverse flips. A leakage audit of the primitive fields flagged 4 affected primitives, which were manually revised.

  • Primitive scoring: Each predicted primitive is compared against the human-annotated gold primitive, with the reference solution used only to recognize equivalent formulations. Each problem belongs to one of three primitive families (Recast, Witness, Argument), and primitives are anchored to a taxonomy of ten primitive types plus a computation_only escape hatch. The rubric combines validity (V), structural identification (σ_gate), and mechanism explanation (σ_mech), with Score = V · σ_gate · (0.6 + 0.4 · σ_mech) and PrimitiveAcc = 1[Score ≥ τ], where τ = 0.8.

  • The four dimensions: Discovery is π(x) → p̂ (find the primitive from the problem alone); Generation is π(x) → ŷ (solve directly, scored by final-answer exact match); Digestion is π(x, y) → p̂ (extract the primitive once a correct solution is given); Execution is π(x, p) → ŷ (solve once the primitive is supplied).

  • Diagnosis: 12 models are evaluated with high reasoning effort where available and a maximum generation budget of 120k tokens per problem. Failures are decomposed by the joint outcomes of Discovery (D) and Execution (E).

  • Post-training experiments: The training corpus contains 709 Mathematics Ph.D. qualifying-examination problems from 1991–2026, each paired with a human-authored proof (median length 118 words), chosen because their open-ended, proof-oriented nature matches Prim and because they generally lack algorithmically verifiable short answers — so the authors use text-supervised post-training rather than outcome-verified methods such as RLVR. They evaluate SFT and OPSD on Qwen-3.5-27B, 9B, and 4B, using the default OPSD configuration, bfloat16 precision, a fixed seed of 42, and NVIDIA RTX 6000 Ada GPUs.

  • Absorb: Rather than supervising the student to output primitives, the teacher is conditioned on the primitive as privileged information while the student generates its own on-policy trajectory. Because naive OPSD causes privilege leakage on complex reasoning problems, Absorb uses reverse KL with one-sided clamping over the teacher's top-K token support, so positive teacher guidance toward better alternatives transfers fully while excessive suppression of student-preferred choices is bounded.

  • Final evaluation: Four benchmarks are used — Prim (Generation and Discovery), the full mathematics subset of HLE-Verified, HMMT25, and Omni-MATH — with pass@4 accuracy reported for HMMT25 and Omni-MATH. The reported average is computed exclusively over Generation scores across all four benchmarks, excluding Discovery.

Why This Matters

The paper shifts attention from whether a model produces a correct answer to whether it possesses the structural understanding that makes the answer possible, and it shows that this distinction has practical consequences for how models should be trained.

Impact on research: It provides a diagnostic vocabulary (Discovery, Generation, Digestion, Execution) and a benchmark for separating structural understanding from procedural execution, and it demonstrates that a diagnosed bottleneck can be targeted directly through post-training design rather than through generic data scaling.

Real-world applications:

  • Mathematical agents and tool use: The observed discovery–execution asymmetry suggests explicitly separating structural search from procedural execution as a design principle for allocating inference-time computation.
  • Scientific domains beyond mathematics: The diagnostic framework is proposed as generalizable to other structurally rigorous fields such as theoretical physics and computational chemistry.
  • Formal verification: The authors propose integrating mathematical primitives with formal verification systems such as Lean, connecting informal structural reasoning with formal proof generation.
  • Post-training pipelines: The bounded override mechanism offers a concrete recipe for transferring privileged teacher information without the regressions seen with SFT and OPSD.

Industry relevance: The results matter directly to anyone fine-tuning reasoning models. The paper documents that SFT and OPSD can degrade Generation accuracy on problems a model already solves (a 6.04-point drop for SFT at 9B, a 4.95-point drop for OPSD at 9B, and a 3.30-point OPSD drop at 27B), while Absorb improves Generation at all three tested scales. That is a practical, measurable argument about how supervision from reference solutions should be transferred.

Future Directions

  1. Generalize the diagnostic framework beyond mathematics to structurally rigorous domains such as theoretical physics and computational chemistry.

  2. Exploit the discovery–execution asymmetry in agent design, for example by explicitly separating structural search from procedural execution to allocate inference-time computation more effectively.

  3. Connect primitives to formal verification, such as integrating mathematical primitives with Lean to bridge informal structural reasoning and formal proof generation.

  4. Refine the transfer of privileged supervision. The paper reports that a substantial fraction of discovery-limited failures remain unrepaired (SFT and OPSD repair roughly 20.4% and 21.1% of D−E+ cases) and that regressions persist, raising the open question of how to transfer guidance even more selectively.

Target Audience

Researchers and practitioners working on LLM reasoning, mathematical problem solving, and post-training methods (SFT, distillation, self-distillation). It is also relevant to evaluation researchers interested in process-oriented rather than outcome-only metrics, and to engineers building mathematical or scientific reasoning agents who need to decide where to spend inference-time compute. Readers interested in AI for Science and in the interface between informal reasoning and formal proof assistants will find the framing useful.

Authors’ abstract

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.

Read the original paper