Research
On the Off-Policy Teacher in On-Policy Distillation
On the Off-Policy Teacher in On-Policy Distillation — Plain-Language Summary Overview Research area: Large language model post-training, specifically on-policy distillation (OPD) and reinforcement lea

- arXiv
- 2609.38360
- Published
- 2026-09-29
- Authors
- Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Blöbaum, Purak Jain
AI summary
On the Off-Policy Teacher in On-Policy Distillation — Plain-Language SummaryOverview
Research area: Large language model post-training, specifically on-policy distillation (OPD) and reinforcement learning with verifiable rewards (RLVR).
Technical level: Advanced. The paper assumes familiarity with knowledge distillation, KL divergences, policy-gradient methods such as GRPO, and next-token entropy diagnostics.
Scope: The paper diagnoses a "teacher-side distribution shift" in OPD — where the teacher must supervise prefixes generated by the student rather than by itself — and proposes SCOUT (Student-COnditioned Updates of the Teacher), a co-training framework that periodically adapts the teacher to student-generated prefixes with outcome-reward RL, evaluated across three math teacher–student pairs and one code setting.
What This Paper Is About
On-policy distillation trains a student model on its own generated trajectories while a stronger teacher supplies dense token-level supervision at every step. This creates an asymmetry the paper calls the off-policy teacher issue: the trajectories are on-policy for the student but off-policy for the teacher, which was optimized to continue from its own prefixes and now must supervise prefixes it would rarely produce itself. The authors show empirically that the teacher's continuation accuracy degrades as student prefixes grow longer and that its next-token entropy stays persistently high on student-generated contexts, then propose training the teacher itself — rather than only filtering which of its signals the student uses — to fix this.
Key Contributions
-
Diagnosis of the off-policy teacher problem. The authors identify teacher-side distribution shift as a distinct challenge in OPD and frame teacher adaptation as a complementary optimization axis to existing work that regulates or restricts a frozen teacher's supervision.
-
The SCOUT method. SCOUT keeps the standard OPD student update unchanged and adds periodic teacher updates: given a student-generated prefix, the teacher samples continuations, receives verifiable outcome rewards, and is updated with RL, with the refreshed teacher synchronized back for later OPD steps.
-
A rigorous multi-setting evaluation framework. The framework spans two task domains (mathematical reasoning and code generation), multiple teacher–student pairs, model scales, model families, and repeated runs with multiple seeds and repeated per-benchmark evaluations.
-
Analysis isolating the mechanism. Controlled comparisons show the gains come from conditioning teacher updates on student-generated prefixes rather than from extra teacher training alone, and that SCOUT is complementary to loss-level OPD improvements such as OPTR.
Main Findings
-
Teacher continuation degrades on longer student prefixes. Using AIME 2025, 4 student responses per question truncated at prefix ratios from 10% to 90%, and 4 independent teacher continuations per prefix, final-answer accuracy declined as prefixes grew longer, while generated continuations stayed well within the 16,384-token budget — ruling out insufficient generation length as the cause.
-
A persistent entropy gap. Teacher entropy on its own prefixes decreased over the trajectory, whereas entropy on student-generated prefixes increased and then remained high, creating a persistent gap that indicates greater teacher uncertainty on student contexts.
-
SCOUT improves OPD across all three math settings. Average accuracy improved by 1.2 to 2.6 points over standard OPD with a frozen teacher across the three teacher–student pairs.
-
Concrete math results. With Qwen3-4B-Instruct-2507 teaching Qwen3-1.7B, SCOUT reached a mean of 51.4 vs. 49.2 for OPD (+2.2). With Qwen3-8B-DAPO as teacher, SCOUT reached 51.6 vs. 49.0 for OPD (+2.6). With Skywork-OR1-Math-7B teaching DeepSeek-R1-Distill-Qwen-1.5B, SCOUT reached 53.6 vs. 52.4 for OPD (+1.2), while GRPO collapsed to 22.6 and Prune-OPD fell to 36.8.
-
Gains transfer to code generation. With Qwen3-4B-Instruct-2507 teaching Qwen3-1.7B, SCOUT improved the code mean from 56.6 to 59.7 (+3.1), whereas Relay-OPD dropped to 12.1 with frequent syntax errors.
-
Improvement is in the teacher, not just in stronger student prefixes. Holding student prefixes fixed and comparing the initial teacher with the teacher after one epoch of training, the trained teacher was more accurate across all prefix lengths. Matched student–teacher checkpoints at initialization, Step 160, and Step 321 showed continuation curves shifting upward over training.
-
Teacher uncertainty is reduced. After SCOUT training, teacher entropy on student prefixes decreased substantially over later response positions, narrowing the gap with entropy on the teacher's own responses.
-
Extra teacher training alone is not enough. The OPD + Teacher GRPO control used the same outcome-reward teacher updates but started rollouts from the original problem. On 4B math it slightly underperformed OPD (48.8 vs. 49.2); on 4B code it reached 57.0 vs. SCOUT's 59.7; on 8B math it reached 50.8 vs. SCOUT's 51.6.
-
Frequent updates are unnecessary. In the Qwen3-8B-DAPO → Qwen3-1.7B setting, update intervals of 1, 5, and 10 student steps gave similar Avg@6 of 51.3, 52.5, and 51.5, while interval 20 fell close to standard OPD. The main experiments use an interval of 10.
-
Additional OPD training does not close the gap. Training SCOUT for one epoch against OPD for three epochs on the same data, with evaluations every 40 training steps, OPD improved briefly then plateaued while SCOUT continued to improve. The authors note this fixed-data comparison does not rule out further gains from fresh data.
-
SCOUT is complementary to loss-level methods. Combined with On-Policy Trust Region (OPTR, the first component of TrOPD), SCOUT added 2.5 points on math (49.8 to 52.3) and 1.2 points on code (59.3 to 60.5) on top of OPTR, and 2.2 (49.2 to 51.4) and 3.1 (56.6 to 59.7) on top of OPD.
-
Baselines are weaker or brittle. Across Table 1, ESR reached 48.2 math / 55.1 code, Prune-OPD 48.7 / 55.8, Relay-OPD 49.3 / 12.1, and GRPO 50.3 / 62.2, compared with SCOUT's 51.4 / 59.7 (the base Qwen3-1.7B student scored 31.9 math / 45.7 code).
Methodology in Plain English
The authors first check whether the problem is real. They take a student model, let it produce complete answers, cut those answers off at various points, and then ask the teacher to finish the reasoning from each cut-off point. They measure whether the final answer is correct and how uncertain the teacher is token-by-token. The result: the further into a student-generated answer the teacher starts, the worse it does.
SCOUT then attacks this directly. During training, the student generates a trajectory as usual and receives dense token-level supervision from the teacher along that trajectory — this part is unchanged from standard OPD. On top of that, every f student steps, the teacher is given a prefix taken from a student answer and asked to generate several continuations. Because the tasks have checkable answers, each continuation can be scored automatically, and the teacher is updated with group-relative policy optimization using those outcome scores. Gradients are applied only to the tokens the teacher itself generated. The updated teacher is then synchronized back to supervise subsequent student updates.
Since the student keeps changing, the prefixes it produces keep changing too, so the teacher is refreshed periodically rather than continuously — the authors find update intervals of 1, 5, or 10 steps behave similarly. Because longer student prefixes are harder, the framework starts teacher adaptation on shorter prefixes and linearly increases the prefix ratio over training, gradually raising the difficulty.
Why This Matters
Impact on research. Most prior work treats the teacher as fixed and changes where or how its supervision is used — shortening rollouts (ESR, prefix distillation), reweighting or masking tokens (IW-OPD, SOD, TIP, TrOPD, LGR), or letting the teacher intervene mid-rollout (Relay-OPD, MOTAB). SCOUT flips the question to whether the teacher itself can be improved for student-generated states, and the complementarity result with OPTR suggests the two families of fixes address different failure modes.
Real-world applications:
- Mathematical reasoning systems: training smaller, cheaper models to solve competition-style problems (AIME, AMC, HMMT, OlympiadBench, MATH-500) by distilling from larger teachers.
- Code generation assistants: improving student models on LiveCodeBench v5, HumanEval+, and MBPP when the teacher and student come from different lineages.
- Post-training pipelines with mismatched teacher–student pairs: situations where the strongest available teacher is from a different family than the deployed student, as in the Skywork-OR1-Math-7B → DeepSeek-R1-Distill-Qwen-1.5B setting.
- Verifiable-reward training loops: settings where correctness can be checked automatically, making outcome-driven teacher updates practical.
Industry relevance. The method reuses the same student-side distillation objective as standard OPD and only touches the teacher periodically, so it is a drop-in addition rather than a new pipeline. The paper also shows the approach is robust in a setting where GRPO collapses and Prune-OPD degrades substantially, which matters for teams running distillation at scale across model sizes and families.
Future Directions
-
Generalizing beyond verifiable rewards. SCOUT relies on outcome rewards from checkable answers; extending student-conditioned teacher adaptation to domains without automatic verification is an open problem.
-
Testing with fresh data. The compute-matched comparison trained OPD for three epochs on the same data, which the authors explicitly note does not rule out further gains from fresh data; how SCOUT behaves under a larger or refreshed data budget remains untested.
-
Scaling to larger models and more domains. The evaluation covers teacher sizes from 4B to 8B and two task domains; whether the same benefit holds at larger scales and in other reasoning domains is not reported.
-
Cheaper or smarter adaptation schedules. The prefix ratio is increased linearly and the teacher is updated every f steps; whether other schedules, alternative teacher objectives, or different synchronization strategies yield further gains is left open.
Target Audience
Researchers and engineers working on LLM post-training, knowledge distillation, or RLVR who already understand policy-gradient objectives and token-level distillation — particularly those building distillation pipelines where the teacher and student differ in scale, family, or training recipe. Readers looking for an accessible introduction to distillation will find the preliminaries section helpful but the analysis sections assume prior background.
Authors’ abstract
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.