Skip to content
AI.info

Research

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Overview Research area: Machine learning — knowledge distillation, on-policy

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
arXiv
2609.36484
Published
2026-09-29
Authors
Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang, Naibo Wang

AI summary

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Overview

Research area: Machine learning — knowledge distillation, on-policy distillation, and reinforcement-learning (RL) post-training of language models.

Technical level: Advanced. The paper combines a training-objective derivation with two propositions, layerwise hidden-state analysis, and multi-seed benchmark evaluation across four base/teacher model pairs.

Scope: The paper proposes RIDE (RL-Induced Direction Extrapolation), a method that trains a student language model by regressing its hidden states toward targets placed beyond an RL-trained teacher along the layerwise representation change that RL produced, and evaluates it on competition mathematics benchmarks.

What This Paper Is About

On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on the student's own sampled trajectories, usually treating the teacher as the final endpoint of learning. Prior work showed students can surpass the teacher by extrapolating an implicit reward in output (logit/log-probability) space, but that signal is attenuated unevenly by the language-model head and is estimated from a single sampled token, which injects noise that extrapolation amplifies. This paper asks whether that same teacher-to-base displacement can instead be measured directly in hidden-state space, before the head, and extrapolated there with a single coefficient.

Key Contributions

  1. Identification of the RL-induced representation residual. The authors define the residual as the layerwise hidden-state difference between an RL-trained teacher and its pre-RL checkpoint, evaluated on identical on-policy prefixes. Both checkpoints are available in the same-initialization distillation setting (student initialized at the pre-RL checkpoint).

  2. The RIDE method. RIDE displaces the representation-matching target beyond the teacher along this residual using a single coefficient λ: h* = h_T + (λ−1)·(h_T − h_B). At λ = 1 the objective recovers OPRD exactly, so any difference from OPRD is attributable solely to the residual term. The authors interpret the regression as maximizing a linear directional reward under a quadratic penalty centered at the teacher.

  3. Empirical demonstration across four base/teacher pairs. RIDE approaches or exceeds its RL-trained teacher on every one of the four pairs (R1-Distill-1.5B, Qwen3-4B, Llama-3.2-3B, Phi-4-mini) and is the only evaluated method whose mean does so, while consistently outperforming OPRD and output-space extrapolation.

  4. Analysis of why representation space matters. The paper characterizes head attenuation (the head's anisotropic singular spectrum) and proves a conditional-variance result showing that output-space extrapolation scales sampled-token noise by (λ−1)², whereas the RIDE gradient is deterministic given the rollout. A four-way direction ablation isolates the residual direction as the source of the gain.

Main Findings

  • RIDE is the only method whose mean exceeds the teacher on all four pairs. On the R1-Distill-1.5B → JustRL-1.5B pair, RIDE reaches Avg@16 of 56.38 versus the teacher's 55.30 and the untouched student's 39.00. On Qwen3-4B → Just-Qwen3-4B, RIDE scores 66.07 versus the teacher's 65.59. On Llama-3.2-3B → Just-Llama-3.2-3B, RIDE scores 13.33 versus the teacher's 13.01. On Phi-4-mini → Just-Phi-4-mini, RIDE scores 18.30 versus the teacher's 17.96.

  • The margin over OPRD isolates residual extrapolation. Because RIDE differs from OPRD only in its target, the margin over OPRD (0.97 to 4.06 points across pairs) measures the benefit of residual extrapolation under a shared protocol. OPRD scores 54.50, 62.01, 10.94, and 17.33 on the four pairs.

  • Margins over the teacher are often within seed noise. On three of the four pairs, RIDE's margin over the teacher is within one across-seed standard deviation (reported in Table 4); on the Llama-3.2-3B pair, the AIME benchmarks lie near the evaluation floor, so evidence there rests mainly on AIMO.

  • Output-space extrapolation (ExOPD) degrades the student. ExOPD falls below its teacher on every pair. On the R1-Distill pair it also falls below the output-space teacher-matching baselines (49.87 versus 52.90 for sampled-token OPD) even though RL moved that teacher 16.3 points from its base. On the three pairs where the teacher-to-base change is small (2.8 to 6.0 points), it falls below the untouched student as well, by 14.1 points on Qwen3-4B. On average RIDE improves over ExOPD by 9.1 points.

  • Head attenuation is anisotropic and weakens the output-space signal. The head's 512 weakest directions carry 79.8% of the teacher-to-base residual's hidden-state energy but only 30.0% of its centered-logit energy. The residual retains only 0.59 of the head gain of an equal-norm isotropic direction. Those weak directions receive 73.0% of the representation-loss gradient but only 48.6% of the output-KL gradient at the head input. The centered head is full rank, with singular values spanning 71.7 to 1.8, so the residual is attenuated by a factor of up to 40 along its dominant directions rather than removed.

  • The student continues past the target along the residual direction. At λ = 1.25, the student's update aligns with the residual at cosine 0.954, its head-visible component tracks the target at cosine 0.895, and its projection onto the residual is 1.69 — above the target coefficient of 1.25. The teacher's head drifts from the base head by only 1.65%, supporting the shared-head assumption.

  • Deterministic versus sampled learning signal. On 9,152 on-policy validation prefixes per method, the conditional variance of the ExOPD advantage rises 12.6× from λ = 1 to λ = 2, with the extrapolation term dominating beyond λ ≈ 1.3, whereas resampling the next token leaves the RIDE per-position gradient norm unchanged. The negative covariance term in the variance decomposition produces a shallow minimum near λ ≈ 1.14, which the authors note matches ExOPD's tolerance of coefficients slightly above one.

  • Coefficient sweep. For RIDE on the R1-Distill-1.5B pair, final Avg@16 rises from 48.2 at λ = 0.5 to 54.3 at λ = 1 and peaks at 55.4 at λ = 1.25; λ = 1.15 and λ = 1.35 give 55.2 and 55.3, so every λ in [1.15, 1.35] improves over OPRD. Beyond λ = 1.35 the degradation is graceful, with λ = 1.5 and λ = 2 finishing at 52.3 and 52.4. All runs with 0.75 ≤ λ ≤ 1.35 have format scores of 94–97%. Displacing beyond the teacher (λ ≥ 1) shortens responses to 5.3–5.7k tokens; displacing toward the base (λ < 1) lengthens them to 6.8–7.4k.

  • ExOPD collapses under extrapolation. ExOPD is best at λ ≤ 1 (52.9 at both λ = 0.75 and λ = 1) and is harmed by every λ > 1: extrapolated runs peak early (by step 150 for λ = 1.25 and 1.5) then decline (λ = 1.25 to 49.9, λ = 2 to 46.2), with the format score falling from 92% at λ = 0.5 to 64% at λ = 2 and responses lengthening to 8,700 tokens.

  • The residual direction matters, not just the displacement magnitude. The ablation on the R1-Distill-1.5B pair holds displacement magnitude fixed and alters only direction. Relative to OPRD (54.50), the random (55.04), reversed (53.90), mismatched-origin (54.75, using Qwen2.5-Math-1.5B-Instruct as the mismatched base), and trajectory-mismatched (55.12) controls all land within one point, while RIDE with the RL-induced residual reaches 56.38.

Methodology in Plain English

The setup requires three models that start from the same checkpoint: a pre-RL base model, an RL-trained teacher derived from it, and a student initialized at the base. Because all three share an architecture, tokenizer, and language-model head, their hidden states can be compared directly on the same text.

On each training step, the student generates its own responses. The frozen teacher and frozen pre-RL base then process those exact same token sequences, and for every layer and token position the authors subtract the base hidden state from the teacher hidden state. That difference is the RL-induced residual, and it is computed rather than sampled, so it is deterministic for a given rollout.

RIDE then sets each training target by extending the teacher's hidden state along that residual, scaled by one coefficient λ. At λ = 1 the target is just the teacher and the method is identical to OPRD. At λ > 1 the target sits beyond the teacher, in the direction RL moved the model. The student is trained with dimension-normalized squared error toward these targets across all layers and the last 2,000 response positions, with the target detached from gradients.

The authors justify measuring the displacement in hidden states with two arguments. First, the language-model head is a linear map, so the output-space extrapolation target is exactly the head projection of the RIDE target — but the head attenuates the residual unevenly, concentrating it in singular directions it amplifies least and constraining no layers below the final one. Second, because output-space extrapolation estimates a log-ratio from one sampled token, extrapolating scales its variance by (λ−1)² no matter how close the student is to the teacher, while the hidden-state gradient has zero conditional variance given the rollout.

Experiments use prompts from DAPO-Math-17K, a budget of 500 training steps shared by all methods, λ = 1.25 with loss scaling by λ⁻², and three independent training seeds (14, 42, 2027). Evaluation reports Avg@16 on AIME 2024, AIME 2025, and AIMO (AMC 2022–2023), where correctness is averaged over 16 sampled responses per problem and then over problems, with Avg. being the unweighted mean of the three benchmark scores. Baselines are sampled-token OPD (top-1), top-16 OPD, OPRD, and ExOPD.

Why This Matters

Impact on research: The paper reframes distillation from an RL teacher as following a direction rather than copying a destination. It shows that the output-space view of OPD as KL-constrained RL with an implicit log-ratio reward has a representation-space counterpart, and it supplies both a theoretical argument (head attenuation, a conditional-variance decomposition) and an ablation showing that the specific RL-induced direction — not merely a displaced or larger target — produces the gain. It also gives a controlled comparison in which ExOPD and RIDE share the same λ and the same pre-RL reference, isolating where the displacement is measured as the only variable.

Real-world applications (as framed by the paper):

  • Distilling an expensive RL run back into its own base model, recovering RL-level capability without paying for the RL run at inference time.
  • Merging domain experts that share a base checkpoint, since the same-initialization requirement holds in that setting.
  • Post-training smaller models (the pairs include 1.5B, 3B, and 4B-scale models) to inherit reasoning gains from larger or more heavily trained counterparts.
  • Curating competition-mathematics capability, the evaluation domain (AIME 2024, AIME 2025, AIMO AMC 2022–2023 problems), where verifiable correct answers make distillation gains measurable.

Industry relevance: Any post-training pipeline that already produces an RL checkpoint alongside its pre-RL base can supply the residual RIDE needs at essentially no extra data cost, since the method only requires two forward passes on the student's own rollouts. Because RIDE keeps the OPRD loss form unchanged aside from the target, it is described as an implementation on top of an existing OPRD pipeline. The reported stability advantage matters operationally: ExOPD's degradation at λ = 1.25 and collapse at λ = 2 (with format scores dropping to 64% and responses lengthening to 8,700 tokens) is the kind of failure that consumes training budget without recovery.

Future Directions

  1. Relaxing the pre-RL checkpoint requirement. RIDE currently requires the pre-RL checkpoint to exist and to share a representation space with the teacher and student. The authors list relaxing these constraints as future work, which would extend the method to teachers whose initialization is unavailable.

  2. Beyond a single global λ. The method uses one global coefficient; layerwise or tokenwise schedules for the extrapolation coefficient are an open question, particularly given that the ablation shows direction, not magnitude, drives the gain.

  3. Beyond mathematical reasoning and one RL recipe. The paper notes that evaluation covers only mathematical reasoning with a single RL recipe, so generalization to other domains and other post-training procedures is untested.

  4. Safety and calibration of students trained on unrealized targets. Because the student is supervised on hidden-state targets that no model has actually produced, and because the target sits beyond the teacher, the authors state that the student's calibration and safety behavior are not guaranteed to match the teacher's and should be evaluated independently before deployment. They also note that any biases or errors RL introduced could in principle be amplified rather than merely copied.

Target Audience

This paper is aimed at machine-learning researchers and engineers working on post-training, distillation, and RL for language models — particularly those who already run on-policy distillation pipelines and have access to both a pre-RL base checkpoint and an RL-trained teacher. It also suits readers interested in mechanistic analysis of how RL changes internal representations, and in why reward extrapolation in output space destabilizes while the same extrapolation applied to hidden states does not. A working familiarity with transformer hidden states, KL divergence, and policy-gradient-style objectives is assumed.

Authors’ abstract

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: https://github.com/xixixixixxxx/RIDE.

Read the original paper