Skip to content
AI.info

Research

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Overview Research area: Machine learning — large language model post-training, specifically multi-teacher on-policy distillation (MOPD) and knowledge transfer from specialized reasoning models. Techni

arXiv
2610.10460
Published
2026-10-07
Authors
Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi, Xiaomin Li, Rohit Jain, Alborz Geramifard

AI summary

Overview

Research area: Machine learning — large language model post-training, specifically multi-teacher on-policy distillation (MOPD) and knowledge transfer from specialized reasoning models.

Technical level: Advanced. The paper works directly in logit space, defines composite targets over a 151,665-token effective vocabulary, and reports gradient-norm and KL diagnostics, though its central idea can be stated simply.

Scope: A controlled comparison of two ways to build the distillation target in multi-teacher on-policy distillation — endpoint supervision versus teacher-relative shift supervision (Δ-MOPD) — across common-domain composition and routed-domain distillation, holding teacher selection fixed.

What This Paper Is About

When several specialized language models are distilled into one student, the standard approach copies each teacher's endpoint policy. But a fine-tuned specialist's output reflects both what post-training changed and preferences it inherited from its own base model. If that base differs from the student's starting checkpoint, endpoint distillation silently transfers the inherited difference too. This paper asks whether transferring each teacher's teacher-minus-base shift, re-anchored at the student's frozen initialization, is a better object to match.

Key Contributions

  1. Identifies target construction as an independent design axis in MOPD, separate from teacher selection, and instantiates it with Δ-MOPD: a composite target built from z_A + Σ (z_Ti − z_Bi) instead of z_A + Σ (z_Ti − z_A). For same-origin teachers (where B_i = A) the construction reduces exactly to endpoint OPD.

  2. Exposes the mechanism that impedes endpoint transfer: the inherited base pull can have a larger gradient than the post-training shift (Γ_base ≈ 2 for Polaris), so base subtraction lowers the teacher-term norm ratio from 5.2:1 to 1.44:1 and reduces target–student KL from 0.152 to 0.031.

  3. Demonstrates stronger signal combination under composition: with three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 Math and 1.95 five-benchmark points; with two teachers it matches endpoint accuracy.

  4. Extends the construction to a second cross-origin teacher via tokenizer projection and to routed-domain distillation, where phased routing shows higher mean performance in both phase orders and a reduced order gap (from 10.50 to 6.42 points), while interleaved routing performs comparably.

Main Findings

  • Inherited pull dominates early learning. Under endpoint supervision, Polaris's base-reference pull has roughly twice the gradient norm of its shift (Γ_base ≈ 2, averaged over checkpoints recorded every 20 steps). The ratio declines during training because only the base-pull term depends on the student, so the inherited component has its greatest relative influence while the specialization is still being acquired.

  • Base subtraction restores balance. With the Nemotron term identical in both arms, the teacher-term norm ratio falls from 5.2 to 1.44 and target–student KL falls from 0.152 to 0.031. The shift composite sits near the equal-norm independence reference (cancellation 0.287 versus 0.293, cosine −0.016), whereas the endpoint composite has low cancellation (0.136), the signature of one dominant term. The ordering appears by step 20 and persists; at M=3, target–anchor KL remains lower for shifts (0.483 versus 0.695).

  • A closer target accelerates cross-origin acquisition. Distilling Polaris alone, Δ-MOPD leads the English-Math macro by 2.87 pp at step 100. Including precursor-scoring overhead, it scores 47.90% at step 100, already exceeding Endpoint's best evaluated 46.54% at step 195; after charging for slower updates, these checkpoints cost 28.0 versus 42.9 allocated H100-hours. At initialization the endpoint target is 4.94 times farther from the student in target–student KL across 16,384 shared prefix states.

  • Two composed teachers: accuracy is matched. In the mechanism run, 44.91% for Δ-MOPD versus 45.14% for the endpoint composite under Avg@K, and 37.73% versus 36.25% under greedy decoding. Δ-MOPD is higher on Science/IF under both decodings, while the endpoint composite is 3.4 pp higher on MATH-500.

  • Three composed teachers: shifts pull ahead. In the scaling run at step 100 (greedy pass@1 macro), Δ-MOPD reaches 33.50 Math versus 29.39 for endpoint composition, and 26.19 versus 24.24 on the five-suite macro. Adding same-origin JustRL raises shift composition by 4.40 Math and 1.78 five-suite points, but endpoint composition by only 0.43 and 0.58. The gain is concentrated in AIME 2025; Science/IF, which receives no in-domain prompts in this run, is 1.29 pp below the endpoint composite.

  • A projected fourth shift adds more. In an unpaired cross-tokenizer extension, a projected fourth shift from a second cross-origin teacher improves all three Math benchmarks, adding 3.32 Math and 1.85 five-benchmark points.

  • Phased routing: consistent gains in both orders. Δ-MOPD ends above its same-order endpoint control in both orders, by 5.68 and 1.60 five-suite points, with both group macros improving in each order. The observed gap between the two phase orders falls from 10.50 points under endpoint supervision to 6.42 points.

  • Interleaved routing: comparable. With 50/25/25 Math/Science/IF mixed in each batch, the five-suite macro differs by only 0.62 pp in favor of Δ-MOPD (+0.91 Math, +0.19 Science/IF), consistent with the idea that shifts help when several teacher signals are combined rather than when each update sees one teacher.

  • Teacher scale context. Under the same greedy 10,000-token protocol, individual teachers score 46.62/28.08/39.21 (Nemotron-1.5B), 49.02/27.50/40.41 (JustRL-1.5B), and 59.29/36.87/50.32 (Polaris-7B) on Math/Sci-IF/five-suite macros.

Methodology in Plain English

The student is trained on its own generated rollouts, guided by frozen teachers. Each teacher is paired with the exact checkpoint that preceded its post-training stage — its "base" — so the researchers can compute a shift: the difference between the teacher's logits and its base's logits, per token.

For any prefix the student visits, the endpoint approach builds a target by adding every teacher's deviation from the student's starting checkpoint. The Δ-MOPD approach instead adds every teacher's deviation from that teacher's own base, then re-anchors the whole thing at the student's frozen initialization. Both targets are softmaxed and matched by minimizing token-level reverse KL.

The comparison is deliberately paired: both arms share initialization, prompts, optimizer, update budget, and evaluation, and all frozen models score identical student prefixes within an arm. Because two of the three teachers (JustRL and Nemotron) are same-origin, their terms are identical in both arms — the only controlled substitution is how Polaris, the cross-origin 7B Math specialist, is represented.

The anchor and student initialization is DeepSeek-R1-Distill-Qwen-1.5B. Runs include a mechanism run (M=2), a scaling run (M=2 versus M=3), an acquisition run (Polaris alone), and routed-domain runs with interleaved and two phased orders. Evaluation uses 1,309 frozen held-out items across AMC 2023, MATH-500, AIME 2025, GPQA-Diamond, and IF-Eval, with five-suite macros computed as unweighted averages of the five benchmark scores. Entries with error bars are mean ± sample standard deviation over five independently trained seeds.

The paper also reports diagnostic metrics that explain why the two targets differ: Γ_base (gradient norm of the base pull relative to the shift), teacher-term norm ratio, centered-shift cosine, cancellation, top-16 sign conflict, and target–student and target–anchor KL.

Why This Matters

Impact on research. The paper reframes a practical question in multi-teacher distillation — "which teachers should supervise this state?" — as separable from "how should their scores define the target?" That gives the field a second axis to vary. It also supplies mechanistic evidence (gradient norms, cancellation, KL geometry) rather than only headline scores, and shows the construction reduces exactly to existing endpoint OPD when teachers share the student's origin, so it is a strict generalization rather than a replacement.

Real-world applications (as directions this work suggests):

  • Merging publicly released specialists, which often come from different base checkpoints and tokenizers, into a single deployable student.
  • Building domain-routed assistants where Math, science, and instruction-following specialists each own a prompt domain.
  • Curriculum-style training pipelines where teachers occupy separate phases, given the reported reduction in order sensitivity.
  • Lower-cost reproduction of strong teacher performance: reaching the endpoint arm's best evaluated Math score with 35% fewer allocated H100-hours.

Industry relevance. Distillation is a standard route from expensive frontier models to deployable smaller ones. The requirement that a teacher's precursor checkpoint be available is a real constraint — the paper lists it as a limitation — but for organizations that control their own fine-tuning pipeline, the precursor exists by construction. The reported compute savings and the reduced sensitivity to phase ordering are directly relevant to training budget planning.

Future Directions

  • Broadening the evaluation scope. The paper studies public reasoning specialists and common-domain composition on Math prompts, with Science/IF measured only through held-out evaluation; broader specialist families and domain-mixed composition are named as natural extensions.
  • Removing the precursor requirement. Δ-MOPD needs each teacher's precursor and a compatible tokenizer or an explicit projection rule, unlike black-box OPD. Whether shift targets can be approximated without the precursor is open.
  • Calibration-aware composition. The policy-ratio form cancels vocabulary normalization constants but does not calibrate differences in logit temperature or confidence across checkpoints, so shift norms reflect both the magnitude of post-training change and the scale on which it is expressed. The authors present calibration-aware composition as complementary to base subtraction.
  • Understanding when shift targets help. The clearest benefits appear when teacher signals are combined at a state, and the phased-routing results are framed as supporting (not conclusive) evidence; the paper notes interleaved routing performs comparably, leaving the boundary conditions of the benefit to be pinned down.

Target Audience

Researchers and engineers working on LLM post-training, distillation, and model merging — particularly those who must combine several publicly released specialists whose base checkpoints differ from the student's initialization. It is also useful for practitioners building routed or phased multi-teacher training pipelines, and for readers interested in mechanistic diagnostics of distillation targets rather than benchmark tables alone. The logit-space formalism and gradient diagnostics assume familiarity with on-policy distillation and reverse KL objectives.

Authors’ abstract

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce $Δ$-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, $Δ$-MOPD exceeds endpoint composition by $4.11$ Math and $1.95$ five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from $10.50$ to $6.42$ points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.

Read the original paper