Skip to content
AI.info

Research

DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation

Overview Research area: Machine learning / large language model training — specifically multi-task on-policy knowledge distillation, where a smaller "student" model learns from a stronger "teacher" mo

DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
arXiv
2609.33711
Published
2026-09-27
Authors
Ao Yu, Weibo Gao, Heng Zhou, Linan Yue, Rui Li, Suyi Liu, Yu Yan, Yizhong Zhang, Qi Liu

AI summary

Overview

Research area: Machine learning / large language model training — specifically multi-task on-policy knowledge distillation, where a smaller "student" model learns from a stronger "teacher" model.

Technical level: Advanced. The paper builds on per-token sampled distillation gradients and introduces a new weighting rule, so familiarity with policy gradients, KL divergence, and language model training is helpful. The core idea itself is intuitive and can be understood without the math.

Scope (one sentence): The paper introduces DuoOPD, a single four-outcome weighting rule for multi-task on-policy distillation that lets the student's correctness decide the direction of feedback while the joint teacher–student outcome decides how the teacher supplies that feedback.

What This Paper Is About

On-policy distillation (OPD) trains a student model on its own generated responses while a stronger teacher provides token-level feedback. The problem is that this feedback is not uniformly reliable across a multi-task mixture: the teacher solves many questions the student misses, but it also fails on questions the student already answers correctly. Standard OPD ignores both outcomes, so on average it pushes down correct student responses and reinforces some failures. The paper's goal is to design one rule — not a set of task-specific settings — that learns from teacher successes while preserving and reinforcing student successes the teacher cannot provide.

Key Contributions

  1. A unified four-outcome rule for multi-task OPD. DuoOPD conditions on each response's outcome rather than its task, so the same settings apply across all tested task mixtures. Student verification sets the direction of feedback (reinforce successes, suppress failures), and the joint teacher–student outcome selects the teacher's context and how feedback weight is allocated.

  2. Consistent gains across two model families and three task mixtures. DuoOPD leads all five baselines (OPD, OPDVR, EOPD, ExOPD, FiRe-OPD) for Qwen3 and Llama, improving mean macro accuracy over OPD by 2.58 and 5.98 percentage points, and also leads on mixtures spanning scientific calculation, instruction following, and code generation.

  3. Two distinctive designs for teacher–student disagreements. When only the teacher succeeds, its verified answer becomes context for scoring the student's failed response; when only the student succeeds, a weight shared within the task reinforces the whole response.

  4. Ablation evidence about where the gains come from. Every one of the four outcome rules contributes, but outcome-based direction alone stays near the gated baseline OPDVR; the two joint-outcome designs for disagreements supply most of the improvement over OPD.

Main Findings

  • All four correctness combinations occur in task-dependent proportions before distillation. For the Llama pair, the student succeeds where the teacher fails on 5.9% of biology responses but 9.8% of physics responses. Across more diverse tasks, both models fail on 1% of chemistry-understanding training responses but 34% of physics-calculation ones.

  • DuoOPD leads all five baselines on the initial biology/chemistry/physics mixture. Qwen3-4B → Qwen3-0.6B macro accuracy: DuoOPD 66.97% versus OPD 64.39%, OPDVR 64.83%, EOPD 65.81%, ExOPD 64.42%, and FiRe-OPD 64.58%. Llama-3.1-8B-Instruct → Llama-3.2-3B-Instruct macro: DuoOPD 73.26% versus OPD 67.28%, OPDVR 68.68%, EOPD 66.49%, ExOPD 66.34%, and FiRe-OPD 67.64%.

  • The margin over the strongest baseline is 1.16 points on Qwen3 and 4.58 points on Llama. The strongest baselines are EOPD on Qwen3 (65.81% macro) and OPDVR on Llama (68.68% macro).

  • Physics shows the largest per-domain gains over OPD. Gains of 3.73 points (Qwen3) and 7.27 points (Llama). Physics is also the domain where only the teacher succeeds most often during training (24.8% and 14.9% of responses).

  • DuoOPD also leads on a second scientific mixture (materials knowledge, chemistry understanding, physics calculation). Macro 63.48%, which is 0.86 points above ExOPD (62.62%), the strongest baseline. Its largest task-level advantage is 3.25 points over the strongest baseline, in physics calculation — where only the teacher succeeds on 26.8% of training responses and both models fail on 34%, whereas both succeed on 89% of chemistry-understanding responses.

  • DuoOPD leads on a heterogeneous mixture (physics knowledge, IFEval instruction following, MBPP code generation). Macro 49.24%, 1.41 points above OPDVR (47.83%), and highest on every individual task. IFEval has the mixture's highest share of responses where only the student succeeds (7.6%), which weight sharing reinforces.

  • It attains the best result on the hardest task in every setting, despite no worst-task objective. Physics for Qwen3 and Llama (62.10 and 69.02, versus at most 60.92 and 63.90 for baselines), physics calculation (43.33 versus 40.08), and MBPP (32.69 versus 31.83).

  • Gains appear both where the teacher succeeds and where it fails. Figure 1 reports added student successes of 1.90 points on Qwen3 and 3.86 on Llama on teacher-solved questions, and 0.67 and 2.12 points on questions the teacher fails.

  • OPD pushes down correct responses. The original-context log-ratio is negative on average even when the student is correct: −1.31 on Qwen3 and −0.22 on Llama when both models succeed, and −1.24 and −0.25 when only the student succeeds. Correct responses make up 64–69% of training responses.

  • Softplus keeps near-zero gated weights from wiping out reinforcement of successes. Gated weights on verified successes average only 0.08–0.09 on DuoOPD's training responses for both model families, versus 0.57–0.65 under Softplus.

  • Every outcome rule contributes. Replacing any single outcome's feedback with OPD's original log-ratio lowers mean macro by 0.76–1.53 points, including the rule for responses where only the student succeeds, which cover only 6.9% of training responses.

  • The two disagreement designs supply most of the gain. Removing the teacher reference costs 1.49 points (65.48% macro); removing within-task weight sharing costs 0.71 (66.26%); removing both costs 2.41, down to 64.56%, close to OPDVR's 64.83%. Outcome-based direction alone explains little of the 2.58-point gain over OPD on Qwen3.

  • Sharing weights within tasks beats sharing across tasks in all three mixtures. 66.97 versus 66.74, 63.48 versus 63.22, and 49.24 versus 48.89, with larger margins (0.26 and 0.35 versus 0.22 points) in the two mixtures whose task scales differ more. Per-task shared weights average 0.56–0.57 on biology/chemistry/physics, but 0.56–0.64 and 0.57–0.64 in the two additional mixtures.

  • Teacher references sharpen correction on the most frequent disagreement. In Qwen3 runs trained with references, this outcome's log-ratio and feedback weight average −1.80 and −2.35, versus −1.35 and −1.91 on trajectories of runs trained without references.

  • The teacher cache is a one-time cost. Caching one teacher response per distinct training prompt takes 30% (Qwen) and 23% (Llama) of an OPD training loop in GPU-time and is reusable across students, seeds, and hyperparameters. Only responses where the teacher alone succeeds (13–23%) are scored with a reference; the scoring and verification stage accounts for 5.9% and 4.7% of OPD's loop. Measured training-loop change is −0.7% (Qwen) and +7.3% (Llama), with student generation accounting for most of the Llama increase.

Methodology in Plain English

Setting up the comparison. The researchers take a frozen teacher model and train a student on a fixed mixture of tasks. At each update, the student generates its own responses to questions, and both models score every generated token under the same prefixes. The difference between the teacher's and student's log-probabilities for a token becomes its feedback weight: positive weights reinforce the sampled token, negative weights suppress it. This ordinary setup is the OPD baseline.

Adding verification. A task-specific verifier — answer matching for multiple choice, instruction checks for instruction following, unit tests for code — labels each student response as correct or incorrect. Before training, the authors also generate and verify one independent teacher response per training question and cache it.

The four-way rule. The student's verified outcome sets the sign of feedback: all tokens of a correct response get positive weight, all tokens of a failed response get negative weight, regardless of what the teacher prefers. The joint outcome then decides the magnitude:

  • Both succeed: positive feedback in the original teacher context, with teacher preferences setting the strength.
  • Both fail: negative feedback in the original teacher context, with teacher preferences adjusting suppression strength.
  • Only the teacher succeeds: the teacher scores the student's failed response with its own verified answer added to the teacher's context, so correction is allocated along the student's trajectory. Only the teacher sees the reference.
  • Only the student succeeds: all tokens of the response get one positive weight shared within that task, computed once per rollout batch with gradients stopped, so a failing teacher's token preferences cannot reshape the reinforcement.

The magnitude function. Unit-temperature Softplus is used so that every valid token gets a nonzero weight. Its first term matches the gated weights used by OPDVR, and the residual supplies nonzero feedback where the gate would otherwise return zero — important because teacher–student log-ratios average below zero even on verified successes.

Experimental protocol. Three task mixtures of increasing heterogeneity are built from SciKnowEval V2 (biology/chemistry/physics; materials/chemistry understanding/physics calculation) plus RLVR-IFeval and MBPP for the third. All methods share data and verifiers within a mixture and run 60 student updates with 48 questions per update (16 per task), one sampled response per question, and a 1,024-token response limit. Evaluation reports avg@8 (eight sampled responses per test question) and a macro average across the three tasks, averaged over three training runs. This is compared against five baselines: OPD, OPDVR, EOPD, ExOPD, and FiRe-OPD.

Why This Matters

Impact on research. The paper argues that verifying only the student is insufficient: OPDVR-style gating fixes the direction of feedback but still uses the teacher the same way whether or not it succeeded. Verified outcomes from both models become a routing signal — the teacher's success supplies a scoring context, and the student's success triggers protected reinforcement. This reframes multi-task distillation as a per-response decision rather than a task-level decision, and the authors note it composes with task-level methods such as teacher selection, task weighting, and worst-task objectives. The paper also reports an ablation showing that direction alone stays near the gated baseline, which is a useful negative result for the field.

Real-world applications (as suggested by the paper's tasks and framing):

  • Deploying one compact model that must serve several different jobs at once — for example a model handling physics questions, instruction following, and code generation — instead of maintaining separate specialists.

  • Scientific question answering, where the mixture spans domain knowledge, conceptual understanding, and quantitative calculation.

  • Instruction-following assistants, where the student may already satisfy most constraint-checkable prompts and distillation must not erode that ability.

  • Code generation, where unit tests provide the binary verifier used in the third mixture.

  • Industry relevance. The practical case for distillation is cost: serving a 0.6B or 3B student instead of a 4B or 8B teacher. The paper reports a one-time teacher cache cost of 30% (Qwen) and 23% (Llama) of an OPD loop that is reusable across students, seeds, and hyperparameters, and a steady-state loop change of −0.7% (Qwen) and +7.3% (Llama), which matters for whether the method is economical to adopt.

Limitations the authors state. DuoOPD relies on a binary verifier; open-ended tasks such as writing would need rules that handle graded or uncertain judgments. Teacher computation could shrink further through shared caches, compressed references, and selective reference use. Each mixture is trained and evaluated on the same tasks, so transfer to held-out tasks is untested.

Future Directions

  • Handle non-binary or uncertain verification. Open-ended generation such as writing currently falls outside the rule's scope, since it depends on a verifier returning 0 or 1.
  • Reduce teacher computation further. The authors point to shared caches, compressed references, and selective use of references as ways to lower the one-time caching cost.
  • Study longer reasoning trajectories. The paper calls for investigating exploration and self-correction over longer chains of reasoning.
  • Test transfer and task scheduling. Whether joint-outcome guidance improves transfer to held-out tasks is open, and the task-level outcome distribution ρ_k is proposed as a potential signal for scheduling tasks during training.

Target Audience

Researchers and engineers working on knowledge distillation for large language models, multi-task training, or post-training pipelines will get the most from this paper, particularly those who already use student-generated rollouts and have access to a programmatic verifier. Practitioners who need to compress a large model into a smaller one while serving several checkable tasks at once will find the setup directly applicable. Readers without a background in policy-gradient-style language model training will need to work through the gradient notation in Sections 2 and 3, though the four-outcome rule itself is described in plain terms and illustrated in Figure 2.

Authors’ abstract

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.

Read the original paper