Skip to content
AI.info

Research

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Overview Research area: Machine learning — specifically knowledge distillation and post-training of large language models, with a focus on training dynamics (gradients, optimizer state, and numerical

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
arXiv
2610.02179
Published
2026-10-01
Authors
Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You

AI summary

Overview

Research area: Machine learning — specifically knowledge distillation and post-training of large language models, with a focus on training dynamics (gradients, optimizer state, and numerical precision).

Technical level: Advanced. The paper assumes familiarity with reverse KL divergence, policy-gradient objectives, Adam/AdamW internals (first and second moments), gradient clipping, and mixed-precision (FP32 master weights vs. BF16) training.

Scope in one sentence: The paper reverse-engineers how multi-teacher on-policy distillation (MOPD) turns supervision choices into parameter changes and task scores, using Qwen3-1.7B with four RL-trained domain teachers (plus SmolLM3-3B diagnostics) to separate the effects of loss averaging, optimizer state, numeric precision, and vocabulary truncation.

What This Paper Is About

Multi-teacher on-policy distillation trains a single student model on the next-token probabilities of several specialized teachers, aiming to inherit all of their strengths at once — but strong teachers do not guarantee that the student actually acquires each teacher's capability. This paper asks why, by tracing the full path from a supervision choice (how losses are averaged, which vocabulary is supervised, which optimizer is used) through the resulting gradients and optimizer updates, to the final per-task scores. The goal is to identify which design decisions genuinely shape capability integration and which are largely absorbed or hidden by the optimizer and by numerical rounding.

Key Contributions

  1. An empirical decomposition of the MOPD pipeline. The authors compare gradients, optimizer updates, and per-task learning curves from identical student parameters, responses, token masks, teacher scores, and optimizer states, isolating the effect of each supervision choice rather than confounding it with data or initialization differences.

  2. A covariance identity linking loss averaging to implicit response weighting. They show that domain token averaging shifts each domain's gradient by Cov_d(T, g) / T̄_d, meaning that correcting unequal domain token shares still leaves longer responses with more weight, whereas domain response averaging removes that term. The identity is verified in FP64 across all 24 batch-domain groups, with a largest relative residual of 1.53 × 10⁻¹⁵.

  3. Evidence that Adam's first moment and BF16 rounding mask teacher differences. Update cosines between averaging rules reach 0.96, and mean cosine across teacher pairs exceeds 0.83 (using the PG and top-64 intersection KL losses), even though raw gradient cosines between averaging rules are only 0.68; and about 97% of FP32 master weights differ from initialization versus only 7–11% of BF16 weights.

  4. A demonstration that better gradient approximation does not reliably produce better task scores. The top-64 intersection KL gradient nearly matches Qwen's full-vocabulary KL gradient (cosine above 0.999), yet the resulting mathematics accuracy is 2.6 points higher than sampled-token PG under response averaging and 2.1 points lower under global token averaging.

Main Findings

  • Loss averaging is implicit response weighting, and its effect is task- and objective-dependent. With equal prompt counts, mathematics supplies about three times as many tokens as instruction following, giving it roughly three times the loss weight under global token averaging; global token averaging assigns mathematics about 44% of the loss weight. Relative to domain response averaging, global token averaging raises mathematics accuracy by 3.0 points and lowers science by 2.1 points under the PG loss. Under top-64, mathematics instead falls by 1.8 points despite its larger weight, while code rises by 1.5 points despite its smaller weight.

  • Averaging rules change gradient direction much more than they change Adam updates. Across six fixed batches, the mean gradient cosine between domain token and domain response averaging is 0.68 (range 0.40 to 0.91), but the one-step update cosine under the same saved Adam state is 0.96. In full 500-step runs, the spread of four-task mean scores across the three averaging rules is 0.28 percentage points under Adam with the PG loss, 0.06 with the top-64 intersection KL loss, and 0.68 points under SGD with the PG loss.

  • Adam's first moment explains most teacher-update alignment. Mean cosine across teacher pairs exceeds 0.83 with the PG and top-64 intersection KL losses when teachers score student responses from their own domains. Resetting the first moment while keeping the second, or using norm-matched SGD, reduces this to near zero. After subtracting the zero-gradient Adam step U(0; H), the mean pairwise cosine of the teacher increments is below 0.01 on domain-specific inputs, rising to about 0.24–0.35 when all teachers score identical shared prefixes.

  • Momentum-free SGD achieves higher four-task mean scores than Adam under every averaging rule. With the PG loss: domain response 39.94 vs. 38.91, domain token 39.46 vs. 38.63, global token 39.26 vs. 38.79.

  • BF16 rounding hides widespread FP32 parameter movement. Under domain response averaging, about 97% of the MOPD student's FP32 master weights differ from initialization, but only 7–9% remain different after BF16 rounding (7–11% including domain token and global token averaging). About one third of parameters account for 90% of total squared change in FP32, versus about 4% in BF16, so rounding makes broad, small changes look sparse.

  • Raising the response cap affects token-averaged objectives most. Comparing 4,096-token and 8,192-token maximum response lengths, relative raw-gradient differences follow global token > domain token > domain response averaging. The 8K-minus-4K mean-score differences at step 100 are +1.05 points (global token), +0.32 (domain token), and −0.83 (domain response).

  • Top-64 intersections nearly recover the full-vocabulary gradient in Qwen. The top-64 intersection retains on average more than 99.9% of total student probability and more than 99.97% of the student's own top-64 mass. Averaging gradients over 1, 16, or 64 samples per fixed prefix raises cosine to the full-vocabulary reference from 0.54 to 0.87 and 0.97.

  • High mean coverage can hide gradient error at the tail, as shown by SmolLM3-3B. Top-64 retains 97.42% of student probability initially and 98.56% after training, yet 8.42% and 3.65% of weighted prefixes respectively retain less than 90% of student probability, and after training relative gradient error spans 2.16–24.54% among diagnostic batches.

  • Supervision breadth barely changes how many BF16 weights move. Starting from identical parameters and Adam state, PG, top-64 intersection KL, teacher-top-64, and full-vocabulary KL each change nearly the same fraction of BF16 weights in one step. Even with a zero current gradient, about 0.14% of BF16 weights change, close to the supervised fractions; resetting only the first moment lowers the fraction, and fully resetting Adam raises it.

  • Gradient fidelity at fixed prefixes does not predict task-level gains. PG is unbiased for the full reverse-KL gradient, and top-64 approximates its direction with cosine above 0.999 in Qwen, yet under response averaging top-64 exceeds PG in mathematics by 2.6 points (four-task mean +0.16), while under global token averaging mathematics falls by 2.1 points (mean +0.21).

Methodology in Plain English

The setup is deliberately controlled. The student is Qwen3-1.7B, and the four teachers (mathematics, code, instruction following, and science) are RL-trained from that same initialization, so all teachers start from the same place as the student. Student responses are generated on-policy at temperature 1 with top-p = 1, using a 2,048-token prompt cap and a 4,096-token maximum response length (raised to 8,192 in the length experiment), with 64 prompts per step split as 16 per domain and target domain weights of 0.25 each.

To isolate the effect of each choice, the authors cache identical batches — the same responses, token masks, teacher scores, student parameters, and Adam moments — and then re-aggregate the same gradients under three averaging rules: global token averaging (all response tokens weighted equally), domain token averaging (fixed domain weights, responses weighted by length), and domain response averaging (fixed domain weights, responses weighted equally). They then apply each aggregated gradient through the saved Adam state, through Adam with the first moment reset, through Adam with both moments reset, and through norm-matched SGD, and measure the resulting held-out KL change. Gradient and update directions are compared by cosine similarity.

Alongside these fixed-batch probes, they run full 500-step training runs across five seeds (42–46) and record per-task learning curves on MATH-500, LiveCodeBench v6, IFBench strict, and GPQA Diamond. Parameter-level analysis tracks FP32 master weights, BF16 model weights, clipping rates, raw gradient norms, optimizer step norms, and cumulative displacement from initialization, counting unique model coordinates and ranking them by squared change. A separate diagnostic subtracts the zero-gradient Adam step to isolate how the current gradient, rather than stored momentum, moves parameters.

Vocabulary-restricted distillation is studied through a top-k intersection loss, which restricts the reverse KL to tokens that are in both the student's top-k and the teacher's top-k sets and renormalizes probabilities within that intersection. The authors motivate the intersection form by an engineering constraint: their SGLang teacher interface accepts one shared list of token IDs for all positions, so querying the student's position-dependent top-k would require fetching the union of candidate sets, which can grow to min(|𝒱|, Lk) scores for a response of length L. Comparisons cover the PG loss, top-16 intersection KL, top-64 intersection KL, a teacher-top-64 reference without renormalization, and full-vocabulary KL.

Why This Matters

Impact on research. The paper reframes MOPD tuning as a joint problem: response weighting, vocabulary supervision, and optimizer settings interact, so improving one axis (e.g., a closer gradient approximation) can fail to improve — or can reverse — outcomes on another. It supplies a concrete diagnostic toolkit (cosine similarity of gradients and updates, FP32 vs. BF16 change fractions, covariance-corrected weighting analysis, and zero-gradient baselines) that other distillation and RL post-training studies can reuse. It also challenges the assumption that unbiased or low-error gradient estimators translate into capability gains, and it questions the default use of momentum in on-policy distillation where the supervised prefixes keep shifting.

Real-world applications (bullets):

  • Building a single deployable assistant model that covers mathematics, coding, instruction following, and science by distilling several specialist teachers, with better-grounded choices about how to weight each domain's responses.
  • Deciding optimizer and precision settings for memory-constrained training pipelines, where BF16 storage can make extensive small parameter changes look like sparse updates.
  • Diagnosing why a multi-teacher or multi-task training run "loses" one teacher's skill even when all teachers are strong, by checking whether long-response domains are absorbing the loss weight.
  • Choosing vocabulary truncation budgets for teacher scoring infrastructure, balancing gradient fidelity against the scoring cost of large candidate sets.

Industry relevance. Companies that post-train open-weight models with distillation from RL-trained specialists must choose averaging rules, top-k budgets, optimizer hyperparameters, and precision formats — often by convention rather than measurement. This paper shows those conventions carry measurable, task-specific consequences, and that the fraction of BF16 weights that visibly move is a poor proxy for how much the model actually learned.

Future Directions

  • Whether a smaller Adam β₁ improves adaptation to shifting teacher signals in MOPD, given that shared momentum aligns updates but may retain directions favored by earlier, now-outdated responses.

  • Whether gradient clipping and Adam's second-moment normalization can explain why losses with similar expected gradients produce different task outcomes.

  • Whether the observed effects hold beyond the studied model families and scales — the authors explicitly note their findings come from limited model families and scales and may differ at larger scale.

  • Extending the covariance-based weighting analysis beyond length to other implicit or explicit response weights (the appendix notes the identity holds for any nonnegative response weight f), including explicit rollout-selection schemes that choose weights by learning signal.

Target Audience

This paper is most useful to researchers and engineers working on post-training and distillation of large language models, particularly those combining multiple specialist or RL-trained teachers. It will also interest optimizer and training-dynamics researchers studying how Adam's first moment, gradient clipping, and mixed-precision arithmetic mediate between gradients and parameter change. Readers need a working understanding of reverse KL distillation, policy-gradient estimators, and Adam internals to follow the quantitative comparisons, though the framing questions — when does balancing domains help, and when does broader vocabulary supervision pay off — are stated accessibly enough for practitioners who want guidance rather than derivations.

Authors’ abstract

Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.

Read the original paper