Skip to content
AI.info

Research

Making Expert Reasoning Learnable with Self-Distillation

Overview Research area: Machine learning / large language model post-training, specifically reasoning improvement from expert human solutions (imitation learning, self-distillation, reasoning distilla

Making Expert Reasoning Learnable with Self-Distillation
arXiv
2602.02405
Published
2026-02-02
Authors
Ethan Mendes, Jungsoo Park, Alan Ritter

AI summary

Overview

Research area: Machine learning / large language model post-training, specifically reasoning improvement from expert human solutions (imitation learning, self-distillation, reasoning distillation).

Technical level: Intermediate. The paper assumes familiarity with supervised fine-tuning, RLVR, chain-of-thought reasoning and KL-divergence-based objectives.

Scope: The paper introduces and evaluates Distribution Aligned Imitation Learning (DAIL), a two-step self-distillation method that converts a small set of expert human solutions into learnable in-distribution reasoning traces for both instruction-tuned and long chain-of-thought reasoning models.

What This Paper Is About

Standard ways of improving LLM reasoning either require the model to sample a correct solution on its own (reinforcement learning with verifiable rewards) or require a stronger teacher model that can already solve the problem. For the hardest problems, neither condition holds, so no training signal can be extracted. This paper proposes learning directly from high-quality expert human solutions instead, and shows that naive imitation of those solutions fails because expert writing is "didactic" — it skips steps that human readers can fill in but that a model cannot. DAIL bridges that gap by rewriting expert solutions into in-distribution reasoning traces and then training with a contrastive objective that discourages imitating shortcuts.

Key Contributions

  1. Distribution Aligned Imitation Learning (DAIL), a two-step self-distillation framework that (a) transforms out-of-distribution expert solutions into detailed, in-distribution reasoning traces and (b) applies a contrastive objective to focus learning on expert insights rather than shortcuts.
  2. A mixed policy rollout generation procedure, in which the student generates tokens that a frozen "teacher" (the same base model conditioned on the ground-truth solution) verifies against a threshold, deferring to the teacher only when verification fails. This preserves the student's natural reasoning flow and self-correction behavior, which direct sampling with the teacher does not.
  3. A contrastive training objective that subtracts a weighted KL divergence to a "negative reference" model — the base model conditioned only on coarse-grained solution waypoints (extracted automatically with a regular expression) — from the positive KL divergence to the teacher, penalizing shortcut imitation.
  4. A new publicly released dataset, e1-proof, of 669 non-verifiable Olympiad proof problems and expert solutions authored by a current International Math Olympiad coach (Evan Chen's website, used with permission), plus the 417-problem e1-verifiable dataset derived from AIME problems 1985–2023 that the target model failed to solve in 32 attempts.
  5. A demonstration that DAIL works on non-verifiable proof problems, which, the authors state, is not possible with standard RLVR without a generative reward model.

Main Findings

  • Sample efficiency: DAIL leverages fewer than 1000 high-quality expert solutions to achieve up to 31% pass@128 gains on Qwen2.5-Instruct and Qwen3. The paper notes expert data can cost upwards of $1,000 per sample, making e1-verifiable (417 problems) and e1-proof (669 problems) intentionally small.
  • RLVR fails in this setting: On the e1-verifiable problems (chosen because Qwen2.5-7B-Instruct never solved them in 32 attempts), GRPO produces a downward shift in the pass@k curve, which the authors attribute to overfitting to rare stochastic successes. GRPO trained on the >40K-example DeepScaleR dataset shows only marginal pass@1 gains on a subset of benchmarks and degrades at larger k.
  • Hint-based RLVR also fails: NuRL + GRPO consistently underperformed GRPO despite training to convergence, which the authors attribute to a learned reliance on hints.
  • Naive imitation degrades performance: Direct SFT on expert solutions and STaR rationalization both reduce performance relative to the untrained model, on all three mathematics benchmarks in Figure 3.
  • Contrastive loss beats NLL: Qwen2.5-7B-Instruct trained with DAIL's contrastive objective outperformed standard NLL across generation settings and metrics; the pass@1 advantage over NLL was larger when training on directly sampled traces than on mixed policy rollouts, and the gap narrowed to roughly 1% at pass@128.
  • Test-time efficiency roughly doubles: Across benchmarks and token budgets, DAIL achieves roughly the same performance as the untrained model with 2× fewer tokens; gains increase as token limits rise from 512 to 4096 and as model size goes from Qwen3-8B to Qwen3-14B.
  • Out-of-domain generalization: On GPQA-Diamond, DAIL matches or outperforms the untrained model across most settings. For Qwen3, pass@128 went from 93.9 to 96.5 at a 1024-token budget, 95.5 to 96.9 at 2048, 93.4 to 96.5 at 4096, and 93.4 to 96.0 at pass@1 across the pooled settings; for Qwen2.5, pass@1 was 55.1 (base) versus 54.7 (DAIL) and pass@128 was 85.9 (base) versus 84.3 (DAIL).
  • Lower training accuracy, better generalization: In Figure 6, NLL baselines achieve the highest training-set performance on e1-verifiable while degrading on unseen benchmarks, and RLVR baselines show high training accuracy despite never seeing expert solutions. DAIL shows lower training performance but superior out-of-distribution generalization, which the authors read as evidence that the contrastive objective suppresses non-robust reasoning.
  • Mixed policy rollouts help reasoning models only: For the non-reasoning Qwen2.5-7B-Instruct, mixed policy rollouts slightly underperformed direct sampling; for Qwen3-8B (think), mixed policy rollouts improved performance, because direct teacher generations frequently referenced the prompt or expert solution.
  • Training efficiency: Data generation is decoupled from the training loop, allowing parallelization, and because the teacher, negative reference and student initially share weights, a LoRA adapter can be enabled or disabled to avoid storing multiple model copies.

Methodology in Plain English

The method starts with a small dataset of (problem, expert solution) pairs. It has two stages.

Stage 1 — rewriting for the student. A frozen copy of the student model is turned into a "teacher" by giving it the problem together with the expert solution. The teacher then produces a new, detailed reasoning trace for that problem. Because the trace comes from the model itself, it is "in-distribution" — it reads like the model's own reasoning, with the skipped steps from the human solution filled in. For instruction-tuned models the authors found that simply prompting the teacher carefully is enough. For long chain-of-thought reasoning models, that direct approach produced traces that reference the expert solution and contain fewer self-corrections, so the authors use mixed policy rollouts instead: token by token, they sample from the student, check whether the teacher assigns that token probability at or above a threshold τ, and only fall back to the teacher's own sample when the check fails.

Stage 2 — training without absorbing bad habits. The rewritten traces can contain "rationalization shortcuts," where the model forced a derivation to reach a known intermediate result rather than truly deriving it. Plain next-token imitation (NLL) would teach the student to reproduce those shortcuts. Instead, the authors define a "negative reference" — the same frozen base model conditioned on only a stripped-down list of key results from the expert solution, which has been automatically extracted with a regular expression. The loss minimizes the KL divergence between the student and the teacher while subtracting a weighted (by γ, between 0 and 1) KL divergence between the student and this negative reference. In effect, tokens that look plausible under the shortcut-only condition but not under full expert guidance are penalized. The authors report training remains stable because the student starts from the same weights as the teacher and negative reference and is anchored by the positive distillation term.

Evaluation uses pass@k with n = 128 samples per problem and k ∈ {1, 2, 4, ..., 128}, on AIME 2024/2025 (60 problems), BeyondAIME (100 problems), IMO-AnswerBench (400 problems) and GPQA-Diamond (198 questions). Models trained are Qwen2.5-7B-Instruct on e1-verifiable and Qwen3-8B (think) and Qwen3-14B (think) on e1-proof. Baselines include SFT on expert solutions, STaR rationalization, sampling at temperatures in {0.6, 0.8, 1.0, 1.2}, GRPO, NuRL, and GRPO on the >40K-example DeepScaleR dataset. For the token-budget experiments the authors use budget forcing by appending an end-think token and prompting for a final answer.

Why This Matters

Impact on research. The paper targets a recognized bottleneck: on problems too hard for a model to solve and too hard for any current frontier model to solve, RLVR yields zero gradient. It argues that expert human solutions are a usable substitute if they are first translated into the model's own distribution, and it is, to the authors' knowledge, the first demonstration of improving long chain-of-thought reasoning while learning directly from expert solutions. It also opens a path for training on non-verifiable problems such as proofs, where correctness-based rewards do not exist.

Real-world applications.

  • Post-training reasoning models on proprietary or licensed expert datasets where collecting solutions is expensive and sample counts are small.
  • Adapting long chain-of-thought models to new supervised fine-tuning datasets — the authors note there is currently no good way to fine-tune a model like Qwen3 on a set of input-output pairs while retaining reasoning ability, short of augmenting it with offline-generated math reasoning.
  • Domains where reward signals are sparse or undefined, such as Olympiad-style proof writing collected in e1-proof.
  • Teaching models to reason in a human-like manner about safety-relevant behavior, such as refusals and privacy policies, which the authors list as a possible application in the impact statement.

Industry relevance. Cost is central: the paper cites expert data at upwards of $1,000 per sample and notes NuRL required roughly 1k GPU hours to converge, versus DAIL's decoupled, parallelizable data generation and single-copy memory footprint enabled by LoRA. Efficiency gains of roughly 2× fewer tokens for equivalent performance translate directly into serving cost, and the result that gains scale with both model size and test-time compute suggests the method fits scaling-oriented training pipelines.

Future Directions

  • Extending beyond mathematics. The authors report preliminary results on non-math domains in Appendix B — a 2×2 maze task and a Rotten Oranges task from the Reasoning Gym benchmark, plus HealthBench, for which they train on 1529 examples from the consensus subset and evaluate on the hard portion — and call for further exploration of these applications.
  • Applying DAIL to safety and alignment data, such as getting models to reason about refusals and privacy policies in a human-like manner, as suggested in the impact statement.
  • Operating at the frontier where no stronger teacher exists. The paper notes distillation's key limitation is its presumption of access to a stronger teacher, and DAIL is positioned as a way to learn from high-value, non-verifiable data instead.
  • Open questions the paper does not fully resolve, including how the method behaves on further task types, the mismatch between the two reported counts for the proof dataset (683 collected versus 669 after deduplication), and the exact behavior at pass@1 on out-of-domain data, where DAIL slightly trails the Qwen2.5 baseline at some inference settings.

Target Audience

Researchers and engineers working on LLM post-training, reasoning model training, reinforcement learning with verifiable rewards, and distillation. It is most useful to readers who already understand supervised fine-tuning and chain-of-thought training and who face the practical problem of extracting training signal from small, expensive, expert-authored datasets — particularly those in domains where correctness cannot be automatically verified. Readers looking for a beginner-level introduction to language model training would need additional background.

Authors’ abstract

Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct solution to be reinforced or the existence of a stronger model able to solve the problem. However, many difficult problems remain intractable for even current frontier models, preventing the extraction of valid training signals. A promising alternative is to leverage high-quality expert human solutions, yet naive imitation of this data fails because it is fundamentally out-of-distribution: expert solutions are typically didactic, containing implicit reasoning gaps intended for human readers rather than computational models. Furthermore, high-quality expert solutions are expensive, necessitating generalizable, sample-efficient training methods. We propose Distribution Aligned Imitation Learning (DAIL), a two-step self-distillation method that bridges the distributional gap by first transforming expert solutions into detailed, in-distribution reasoning traces and then applying a contrastive objective to focus learning on expert insights and methodologies. We find that DAIL can leverage fewer than 1000 high-quality expert solutions to achieve up to 31% pass@128 gains on Qwen2.5-Instruct and Qwen3, double reasoning efficiency, and enable out-of-domain generalization.

Read the original paper