Skip to content
AI.info

Research

Scaling Properties of Same-Family On-Policy Distillation

Overview Research area: Machine learning — large language model post-training, reinforcement learning, and knowledge distillation (specifically on-policy distillation for mathematical reasoning). Tech

Scaling Properties of Same-Family On-Policy Distillation
arXiv
2609.32722
Published
2026-09-26
Authors
Yuntai Bao, Qinfeng Li, Guoqing Jiang, Liwei Chen, Zhiheng Qin, Xuanping Li, Wenqi Zhang, Xuhong Zhang

AI summary

Overview

Research area: Machine learning — large language model post-training, reinforcement learning, and knowledge distillation (specifically on-policy distillation for mathematical reasoning).

Technical level: Intermediate. The paper assumes familiarity with policy-gradient RL, KL divergence, and distillation objectives, but its central claims are stated in terms of accuracy-vs-training-progress curves and power laws that are straightforward to follow.

Scope: This paper fits power laws that predict the outcome of on-policy distillation (OPD) from student scale, teacher scale, and measured teacher accuracy, using 25 teacher–student pairs spanning Qwen2.5 models from 0.5B to 14B parameters.

What This Paper Is About

Reinforcement learning gives language models strong reasoning ability, but retraining every model size with RL is expensive. On-policy distillation offers a shortcut: let an RL-trained "teacher" score the rollouts of a "student" model, so expertise learned once can be copied to other model sizes. The authors ask whether the outcome of such a distillation run can be predicted in advance from the sizes and scores of the teacher and student, rather than discovered only after training.

Key Contributions

  1. A two-phase description of OPD training dynamics. The authors show that every one of their 25 OPD runs begins with a regular "useful-transfer regime" where held-out accuracy (the gold score, G) rises approximately linearly with training progress d, defined as the square root of token-level reverse KL divergence from the student's initialization. After a "transfer endpoint," dynamics become noisy and split into attenuated improvement, saturation, and regression.

  2. Joint power laws for peak gold score and useful-transfer rate. They fit laws expressing peak gold score and initial slope as multiplicative power laws in normalized student size, an "effective" teacher size capped at the student's size, and the effective teacher's remaining error. Laws are fit independently for Vanilla-OPD and Delta-OPD.

  3. Validation by extrapolation and a controlled out-of-sample test. Withholding the largest model scales or the largest student/teacher scales, the joint law extrapolates to held-out scales within 0.7 accuracy points for Vanilla-OPD and 0.4 for Delta-OPD. A separate controlled comparison distills a 7B student from an intermediate 3B teacher checkpoint scoring 66.0 (slightly above the 1.5B RL endpoint's 63.6) and shows only the joint law orders the results correctly.

  4. Scaling effects of three design choices. Delta-OPD (which rewards the teacher's RL-induced policy shift) beats Vanilla-OPD on matched-KL slope in 15 of 17 shared pairs; off-policy cold start harms weak-to-strong transfer; and bootstrapping weak-to-strong OPD through intermediate model sizes does not improve on direct transfer from the smallest expert.

Main Findings

  • Regular initial regime with heterogeneous tails. Linear fits of gold score against d over the first 30 checkpoints reach R² ∈ [0.932, 0.988] with RMSE ∈ [0.0026, 0.0129] accuracy units, and fitted slopes span [0.175, 0.721]. After the transfer endpoint, the tails mix attenuated improvement, near-saturation, and regression with no shared functional form.

  • Square-root KL is the right coordinate. Along a smooth training path, gold score changes to first order in the parameter perturbation while KL changes to second order, giving G(d) = G(0) + m·d + O(d²) with a slope set by the Fisher-normalized alignment of the update direction with the gold-score gradient.

  • Smaller teachers sometimes hurt larger students. Peak gold score increases with teacher scale for the bigger students — from 72.7% to 81.9% for the 7B student and from 77.7% to 87.3% for the 14B student — but the 0.5B student peaks at 40.8% with the 3B teacher, falling to 39.0% under the 7B teacher and 37.7% under the 14B teacher. The fixed-3B-student series peaks with the 3B teacher rather than the 7B teacher.

  • Peak location is not monotone. For the 7B student, the locations of the gold-score maximum are 0.274, 0.282, 0.308, 0.323, and 0.323 as teacher size increases. The authors therefore track peak value, initial rate, and transfer extent rather than fitting a law to the location.

  • Late regression is implicit-reward overoptimization. Wherever gold score regresses, the logged teacher-induced proxy score P keeps rising, which distinguishes proxy overoptimization from simple over-imitation.

  • Peak error nearly tracks teacher error. With ζ ≈ 1 for both methods, the fitted joint peak law gives A = 0.97, α = 0.30, β = −0.27, ζ = 0.95 for Vanilla-OPD and A = 1.02, α = 0.34, β = −0.33, ζ = 1.01 for Delta-OPD. The negative β means that at a matched score, the smaller teacher transfers better.

  • Teacher score is not the same as teacher value. When a 7B student is distilled from an intermediate 3B checkpoint scoring 66.0 against a 1.5B endpoint scoring 63.6, the observed peaks are 73.8 and 77.5 respectively. The joint law predicts the first within 0.3 points (74.1) and gets the ordering right; the scale-only and score-only laws both favor the larger, higher-scoring teacher (79.5 and 76.3 respectively).

  • Transfer slows sharply with teacher error. The rate law gives ξ = 1.9 for Vanilla-OPD and 2.01 for Delta-OPD (B = 0.12, γ = 0.19, δ = −0.61 for Vanilla; B = 0.13, γ = 0.20, δ = −0.73 for Delta), meaning larger teachers transfer more slowly at a matched score. Rate fits are noisier than peak fits, with log-space R² of 0.59 and 0.76.

  • Transfer extent is roughly scale-free. Departures and censored lower bounds span d of 0.20–0.36 for Vanilla-OPD and 0.27–0.34 for Delta-OPD, within a factor of two across the grid, with medians of 0.30 and 0.29; a censored accelerated-failure-time fit finds only weak scale dependence.

  • Delta-OPD helps most in weak-to-strong pairs. First-40 Delta-OPD lines reach R² ∈ [0.916, 0.987] and RMSE ∈ [0.0049, 0.0116]. Slopes exceed Vanilla-OPD in 15 of 17 shared cells (the exceptions are the two smallest students under the 0.5B teacher), and Delta-OPD attains the larger baseline-normalized peak gain and the larger absolute peak in 12 cells each, mostly in weak-to-strong pairs (nine of ten). The advantage shrinks with teacher scale: the 0.5B teacher gives two to four extra points of absolute peak, while larger teachers stay within about one point.

  • Off-policy cold start harms weak-to-strong OPD. Pure OPD attains the best peak gold score in every cell. One epoch of SFT on the 0.5B expert's rollouts costs 6.6, 15.7, and 19.4 points for the 3B, 7B, and 14B students respectively, and pins every student near the teacher's own score, erasing 32 points of the 14B student's initialization. Pure OffPD underperforms pure OPD in every cell, by 28.4 points in the extreme weak-to-strong cell and by less than one point in same-base cells.

  • Bootstrapping does not beat direct transfer. Every bootstrapped chain peaks below direct OPD from the 0.5B expert at the same scale: 60.4 versus 61.4 at 3B, 70.0–71.3 versus 72.7 at 7B, and 76.8 versus 77.7 at 14B. The three Delta-OPD chains behave alike (64.5 versus 65.3 at 3B, 74.0 versus 75.0 at 7B, 79.4 versus 79.9 at 14B). A 1.5B OPD product scoring 53.0% still produces a 3B student that peaks below the 0.5B RL expert's direct student, even though the expert itself scores only 39.8%. Direct RL on the student remains above every weak-teacher variant.

Methodology in Plain English

The study uses Qwen2.5 Base models (not instruction-tuned) at 0.5B, 1.5B, 3B, 7B, and 14B parameters. All models first undergo supervised fine-tuning on a subset of Dolci-SFT to establish basic instruction-following under a chat template. Teachers are then produced by GRPO reinforcement learning on the mixed GSM8K and MATH training split of 14.8K examples and evaluated on the corresponding test split of 6.3K examples. Distillation uses the same training prompts, runs for at most ten epochs (580 updates), and is implemented in the verl framework, with periodic held-out evaluation. The design covers 25 teacher–student combinations spanning weak-to-strong, same-base, and strong-to-weak setups.

To track training progress, the authors measure a token-mean reverse KL divergence between the current student and its fixed SFT reference using the nonnegative, unbiased k₃ estimator, and take its square root as the coordinate d. They then read each run as a curve of gold score against d, fit a straight line to the first 30 checkpoints, and define the transfer endpoint as the last point before the curve falls below the 95% predictive band of that fit for three consecutive checkpoints.

For predictions, parameter counts are normalized by one billion, and teacher size is capped at student size (Ñ_T^eff = min(N_T, N_S)/(1B)) to approximate the saturation and reversal observed above student scale. The peak and slope are then fit as multiplicative power laws in student size, capped teacher size, and the capped teacher's remaining error. Because teacher size and teacher score are nearly collinear on the RL-endpoint grid (r = −0.9999), the authors add bootstrapped-chain teachers — which are themselves OPD products sitting below the size trend — to break the collinearity (lowering it to −0.95 and −0.97) and identify the teacher exponents.

Why This Matters

This work turns OPD from a run-and-see procedure into a roughly predictable one. If peak student accuracy can be estimated from student size, teacher size, and measured teacher score before training begins, then RL expertise can be trained once at small scale and transferred across a model family with a known accuracy budget. It also corrects a natural but wrong assumption: a teacher's own score alone does not determine its supervision value, and a larger teacher is not always the better teacher.

Real-world applications implied by the paper's results:

  • Post-training pipeline planning. Estimating expected peak accuracy before committing compute to a distillation run, and choosing teacher–student pairs from a fitted law rather than by trial.
  • Amortizing RL cost across a model family. Training one small RL expert and transferring it predictably upward to 3B, 7B, and 14B students, instead of repeating RL at every scale.
  • Reasoning-model distillation for math. The experimental setting is GSM8K- and MATH-style math reasoning, so the laws apply directly to building math reasoning capability in smaller deployable models.
  • Avoiding harmful pipeline steps. The results show that an off-policy SFT cold start and bootstrapped multi-stage chains both reduce final accuracy in weak-to-strong settings, giving practitioners concrete reasons to skip them.

Industry relevance: Frontier and mid-tier labs routinely run distillation in post-training pipelines and face exactly the question this paper addresses — which teacher, which student, and how long to train. The finding that pure OPD beats both off-policy cold start and pure off-policy distillation in every tested cell, and that bootstrapping does not help, has direct cost and quality implications for anyone scheduling RL and distillation jobs.

Future Directions

  • Does the square-root-KL linearity hold outside same-family, math-reasoning settings? The study is deliberately restricted to Qwen2.5 same-family pairs and to math reasoning on GSM8K and MATH; the authors note their parameter-count scaling variable is confounded with data scale and training compute, so the laws should be read at compute-optimal settings.
  • Can the transfer extent be given a law? The authors report that extent resists a comparable power law and is better read as an approximately scale-free KL budget, with Delta-OPD exponents not identifiable in their censored fit. Explaining why the useful-transfer budget is roughly constant remains open.
  • Why does bootstrapping fail here when it helped in prior weak-to-strong work? The paper cites the "feature-slots" picture as a candidate explanation — each OPD stage bottlenecks features inherited from the RL expert — but does not test it directly.
  • Do the smallest students' reversals have structure? The authors

Authors’ abstract

*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.

Read the original paper