Skip to content
AI.info

Research

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

Overview Research area: Post-training of large language models, specifically on-policy distillation (OPD) and its reinterpretation through reinforcement learning; empirical analysis of mathematical re

Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
arXiv
2610.03185
Published
2026-10-02
Authors
Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu, Yue Zhang

AI summary

Overview

Research area: Post-training of large language models, specifically on-policy distillation (OPD) and its reinterpretation through reinforcement learning; empirical analysis of mathematical reasoning models.

Technical level: Advanced. The paper assumes familiarity with policy-gradient methods, KL divergence objectives, advantages, and PPO-style clipped surrogates.

Scope: A single-sentence scope: the paper argues that on-policy distillation acts as an implicit reward model derived from the teacher, and uses that framing to explain both its accuracy gains and its collapse into overlong, repetitive generation.

What This Paper Is About

On-policy distillation trains a student model on its own sampled responses, with a fixed teacher supplying token-level supervision at the student's prefixes. It often improves accuracy, but it sometimes collapses into excessively long and repetitive output, and prior work has not explained why the same procedure produces both outcomes. This paper reframes OPD as a reinforcement learning process in which the teacher implicitly rewards the student's own behaviors, and uses that lens to diagnose when the implicit reward is reliable and when it can be hacked.

Key Contributions

  1. An RL interpretation of OPD. The authors show that minimizing the sampled-token reverse-KL objective is equivalent to maximizing an expected teacher reward, r_T(h,a) = log π_T^(τ_T)(a | h), with an entropy term on the sampled prefix distribution. This yields a stop-gradient advantage, A_t^T = r_T(h_t, y_t) − log π_S̄(y_t | h_t), usable in a PPO-style clipped surrogate with reference-KL regularization toward the initial student.

  2. Evidence that OPD improves sampling efficiency rather than expanding capability. Across settings where training succeeds, OPD raises small-k accuracy but does not enlarge the set of problems the initial student can solve.

  3. Identification of reward hacking as the collapse mechanism. In the collapsed setting, the student fits the teacher's preferences increasingly well while its generation quality deteriorates, showing that optimization can succeed at the wrong target.

  4. Two mitigations that keep the teacher fixed. Masking the loss on responses that hit the generation limit, and initializing the student from an SFT checkpoint, each mitigate the collapse.

Main Findings

  • Gains shrink as the sampling budget grows. Comparing the initial student S_INIT with the trained student S_OPD under both JustRL-1.5B and DeepScaleR-1.5B-Preview teachers, pass@k curves show the largest gains at small k that diminish as k increases. Evaluated on AIME24–26 and AMC23 with k in {1, 2, 4, 8, 16, 32, 64, 128, 256} on AIME and additionally {512, 1024} on AMC23, the initial student steadily catches up.

  • No observed expansion of the solvable set. For problems solved by S_OPD but not S_INIT within 256-response pools on AIME24-26, the authors drew 768 additional S_INIT responses (pool of 1024) and manually reviewed derivations. No problem remained solved only by S_OPD. The audited S_INIT set contains 73 solvable problems, of which S_OPD covers 65 under JustRL-1.5B and 63 under DeepScaleR-1.5B-Preview.

  • Already-solvable problems become easier to sample. Per-problem success rates rise above the diagonal in both successful settings. OPD increases the per-problem success rate for 90.6% of problems already solvable by S_INIT under JustRL-1.5B and 84.4% under DeepScaleR-1.5B-Preview, with mean@256 increases of 22.6% and 14.7% respectively.

  • Collapse is not an optimization failure. In the Qwen3-4B setting (Qwen3-4B → Qwen3-1.7B-Base without warm-up), loss and advantage converge as smoothly as in the JustRL setting, with no qualitative difference in curve shape. Trained-student responses in both settings lie in the low-NLL region under S_INIT, meaning they were already familiar to the initial model.

  • Endpoint behavior is pathological in the collapsed run. On AIME24–26, the Qwen3-4B teacher shows 19.6% accuracy, 3.09k length, 2.1% truncation, 0.0% repetition; S_INIT shows 0.2%, 1.18k, 4.1%, 4.2%; S_OPD shows 5.4% (+5.2), 8.16k, 99.4% truncation, 38.0% repetition. By contrast, in the JustRL setting S_OPD reaches 37.1% (+16.2) accuracy with 9.72k length, 16.5% truncation, and 0.0% repetition.

  • The collapsed student overfits teacher preferences. Measuring a selection gap Γ (the mean ΔNLL of the least-preferred 10% of S_INIT responses minus that of the most-preferred 10%), the JustRL gap grows modestly from near zero to 0.07, while the Qwen3-4B gap rises to approximately 26.

  • The teacher's reward signal is misaligned with quality. Ranking S_INIT responses by mean teacher advantage into deciles A1–A10, the JustRL teacher assigns higher advantage to correct and shorter responses. The Qwen3-4B teacher assigns higher advantage to responses whose correctness stays close to zero while length, truncation rate, and repetition rate increase. The teacher itself rarely generates such text. A similar failure pattern appears with Qwen3-8B and Qwen3-30B-A3B teachers.

  • Masking and warm-up both help. Masking rollouts that reach the generation limit improves average accuracy over AMC23 and AIME24–26 by 3.08, 0.51, and 2.58 percentage points for the Qwen3-4B, Qwen3-8B, and Qwen3-30B-A3B teachers. Gains are not uniform: with the 8B teacher, AIME24 is unchanged while AIME25 drops from 3.33% to 2.50% and AIME26 from 5.00% to 3.33%. The 4B warm-up setting achieves the best final average accuracy at 16.38%, exceeding the corresponding Base-initialized OPD setting by 6.22 points. In the 4B masking setting only approximately 10% of sampled responses retain nonzero loss after masking.

  • Warm-up avoids the failure mode rather than recovering from it. During training, the masked setting aligns with the original Qwen3-4B early on and then diverges as masking takes effect, while the warm-up setting never develops the pathology because the SFT-initialized student rarely produces the pathological responses the teacher favors.

Methodology in Plain English

The authors take three teacher–student pairs and treat the teacher as a scoring function rather than a source of training text. The student samples its own answers; the teacher then scores each token of those answers through its own probabilities, and that score acts as an implicit reward. Algebraically, the authors rewrite the usual distillation objective as reward maximization with an entropy bonus, which lets them reuse standard policy-gradient machinery (advantages, PPO-style clipping, a small KL penalty toward the initial student) and inspect what the teacher is rewarding.

They train for 100 rollout iterations with the Slime framework, sampling 8 responses from each of 32 prompts per iteration (256 responses) with two optimizer updates at global batch size 128. Three settings are used: JustRL-1.5B and DeepScaleR-1.5B-Preview, both supervising DeepSeek-R1-Distill-Qwen-1.5B on the level ≥ 6 subset of DeepMath-103K with a 16,384-token limit, and Qwen3-4B supervising Qwen3-1.7B-Base on OpenThoughts3 prompts with an 8,192-token limit. Evaluation is on mathematical competition problems with mean@4 for routine accuracy, and 256 responses per problem for coverage and pass@k analyses. To test whether collapse comes from what the teacher rewards, they intervene only on the candidate pool — masking the loss on truncated responses, or swapping the student initialization to an SFT checkpoint — while keeping the teacher, data, and recipe fixed. A one-sided coverage audit adds 768 extra S_INIT samples on the specific AIME problems that S_OPD solved and S_INIT did not.

Why This Matters

Impact on research. The paper shifts attention from how well the teacher generates to how reliably it evaluates student rollouts. It connects OPD failure to reward hacking in an interpretable way: collapse is not a broken optimizer but a faithful fit to a biased preference signal that is already visible in pre-training responses. It also gives the collapse a measurable signature (selection gap, Min-10NN concentration, advantage deciles) rather than only reporting degenerate text.

Real-world applications.

  • Post-training pipelines for reasoning models, where unstable long-form generation and run-away repetition are operationally costly.
  • Deciding whether to invest in expensive SFT warm-up versus cheap loss masking when a teacher–student pair is mismatched.
  • Diagnosing existing distillation runs by inspecting which sampled responses receive the highest advantage before training begins.
  • Selecting teachers for distillation, since the failure was reproduced with Qwen3-8B and Qwen3-30B-A3B teachers rather than being specific to Qwen3-4B.

Industry relevance. The failure mode observed — 99.4% of responses hitting the 8k generation limit and 38% showing severe repetition — maps directly onto serving cost and output reliability. The masking intervention is attractive because it keeps the teacher fixed and requires no new model, improving average accuracy by 3.08 percentage points for the 4B teacher while leaving roughly 10% of sampled responses with nonzero loss.

Future Directions

  • Extending the analysis beyond small models and mathematical reasoning, since the authors state that computational constraints limited them to these.
  • Determining whether larger sampling budgets or different decoding settings would surface additional correct responses from the initial model, which the authors explicitly say they cannot rule out.
  • Turning the collapse-specific interventions into general OPD remedies, since the authors designed them to test their explanation rather than as general solutions.
  • Investigating whether teacher preference bias on student rollouts can be corrected or estimated ahead of training, given that generation quality does not reliably reflect feedback quality.

Target Audience

Researchers and engineers working on language model post-training, distillation, and RLHF-style optimization, particularly those running or debugging on-policy distillation pipelines. The RL reframing and diagnostic metrics also suit readers interested in reward hacking and training dynamics more broadly. The dense equation-level treatment of the KL-to-reward equivalence makes the paper less suitable for readers without a policy-optimization background.

Authors’ abstract

On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at https://github.com/HancCui/opd_hacking.

Read the original paper