Research
What Does Privileged Information Add to On-Policy Self-Distillation?
Overview Research area: Natural Language Processing / large language model post-training, specifically on-policy self-distillation (OPSD) and learning with privileged information. Technical level: Int

- arXiv
- 2609.20612
- Published
- 2026-09-17
- Authors
- XiuYu Zhang, Wei Chow, Junfeng Fang, Xingyu Zhu, Zhenkai Liang, Tat-Seng Chua
AI summary
Overview
Research area: Natural Language Processing / large language model post-training, specifically on-policy self-distillation (OPSD) and learning with privileged information.
Technical level: Intermediate (the high-level argument is accessible; the loss formulation, profiling statistics and checkpoint-matching analyses are advanced).
Scope: A controlled empirical study that isolates how much a privileged reference (an answer or worked solution shown only to the teacher) actually adds to on-policy self-distillation, using a purpose-built answer-matched dataset and two model families.
What This Paper Is About
On-policy self-distillation improves a language model by letting a frozen copy of the same model act as teacher, with the teacher given extra information such as the answer or a worked solution that the student never sees. Because such training improves over the base model on its own, it is unclear how much of that gain comes from the privileged reference and how much comes from distillation itself. The authors build a dataset where six different reasoning views share one verified answer, then compare each view against a reference-free control to measure exactly what the reference contributes.
Key Contributions
- AMPLE-Math (Answer-Matched Privileged Levels of Explanation): a reusable supervision suite of 5,319 mathematical problems drawn from OpenThoughts-114k, each paired with six answer-matched reasoning views (Answer Only, Gist, Key Points, Clean Solution, Summary, Full Trace) plus audits and frozen splits.
- Controlled comparisons in two model families (Qwen3-1.7B and SmolLM3-3B) that separate the benefit of cross-mode distillation from the added value of a reference, following each reference across checkpoints rather than reporting a single gain over the base.
- Evidence that changing student training trajectories can reverse transfer: with the same teacher, problems, references and evaluation, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families.
- Teacher profiling plus matched loss interventions showing that changing token-level supervision (marker loss, loss window, reference content, prompt templates) often leaves student behavior largely unchanged, and that stopping time explains most of an apparent gain from broader loss coverage.
Main Findings
- Much of the gain is reference-free: With a thinking-enabled teacher supervising direct-response rollouts and thinking-enabled evaluation, the reference-free student improves both in domain and on external benchmarks, and even a wrong-answer control improves over the base.
- More complete solutions do not reliably help more: At step 100 in Qwen, Answer Only and Full Trace each sit about 0.6 points above the reference-free student, with both intervals including zero. The strongest additional-benefit evidence is a polished Clean Solution advantage of 1.30 [0.20, 2.41] points over reference-free training across three seeds, which excludes zero before adjustment but does not survive Holm correction across the six views.
- The reference helps more clearly in SmolLM3-3B: Full Trace adds 2.0 points over reference-free training at step 50 with thinking enabled. By step 100 all three configurations fall below the base, but Full Trace loses less than the reference-free and Answer Only students.
- Gains concentrate on partially solvable problems: The largest gains, with and without privileged information, come on problems the base never solved in four direct-response attempts but solves 62% of the time with thinking enabled.
- Rankings flip with inference mode: At step 100, Qwen's Answer Only and reference-free students exceed Full Trace by 5.32 and 10.81 points under direct-response evaluation. SmolLM3's step-50 Full Trace advantage of +2.00 points with thinking enabled becomes -10.63 when answering directly, and grading only the first 4K tokens reverses Qwen's ordering.
- Training trajectory reverses transfer: Under thinking-enabled training rollouts, step-50 gains over the base become losses for Answer Only, Clean Solution and Full Trace, a penalty that persists across checkpoints and reproduces in SmolLM3 and in Qwen's external benchmarks.
- Correctness alignment depends on the responses being scored: For Full Trace, correct responses receive larger average teacher-student log-probability shifts than incorrect ones on direct-response prefixes, and the ordering reverses on thinking-enabled prefixes. On direct-response-trained students' responses, alignment shrinks for five views while Full Trace retains most of its initial value.
- Suppressed correction markers are not the differentiator: Every view lowers sampled reconsideration-marker probability on average in both prefix modes, so this shared sign does not distinguish configurations with opposite outcomes. Excluding or downweighting loss at these markers produces no accuracy gain and leaves marker use almost unchanged.
- Diagnostic disagreement sits at the start of the response: The first quarter of the retained span carries disproportionate full-vocabulary KL in both prefix modes.
- Reference content still matters: Replacing Clean Solution or Key Points with a length-matched view from another problem including that problem's answer lowers thinking-enabled accuracy by about two points (-1.95 [-3.78, -0.20] and -2.15 [-3.97, -0.39] versus the matched configuration), so reference-free gains do not make content irrelevant.
- Truncating a trace changes which mode benefits: Cutting Full Trace to the length of Clean Solution, Key Points or Summary brings no clear thinking-enabled gain, yet improves direct-response accuracy over matched Full Trace by roughly 8 to 11 points.
- Most of the loss-window gain is stopping time: First-4K and Distributed-1K exceed the development-selected Early-1K by about four points but stop at step 25 instead of step 50; at the matched step 25 the advantage falls below one point, and their intervals include zero at every matched checkpoint.
Methodology in Plain English
The authors start by building a dataset in which answer information is held constant while the form of the reference is varied. Each of 5,319 problems from OpenThoughts-114k is paired with six views that end in the same canonical answer section but differ in reasoning density, from Answer Only to the complete 4,916-token Full Trace. Three views (Gist, Key Points, Summary) are generated from the intact source trace by a larger model and audited for fidelity; Answer Only is deterministic, and Clean Solution and Full Trace are source-derived. Problem difficulty is annotated with 8-16 problem-only samples per question from three evaluator models and a binomial Rasch model, and the frozen base model's own direct-response outcomes define three success bands used to build train (1,536), development (192) and test (384) splits.
Training uses LoRA adapters (r=64, alpha=128) on Qwen3-1.7B for 100 optimizer steps, with SmolLM3-3B providing the cross-family test and its own rebuilt training split. The teacher is always the frozen, thinking-enabled backbone; the primary configuration trains on direct-response rollouts (capped at 1,024 tokens) and evaluates with thinking enabled, an asymmetry that remains even without any reference. The objective is the generalized forward KL between teacher and student next-token distributions with each vocabulary term clipped at 0.05, and because the teacher and student start from the same weights, zero loss would occur without privileged information supplying the initial discrepancy.
Because improvement over the base alone cannot separate the two sources of gain, every reference view is compared against a reference-free control under identical settings, and comparisons are made at matched checkpoints. To explain the results, the authors profile token-level supervision on 512 problems, measuring correctness alignment (do correct responses get larger shifts than incorrect ones), correction pressure at markers such as "wait" and "check", and where KL is concentrated within the scored span. They then test candidate fixes with interventions that hold everything else fixed: swapping in other-problem references, truncating traces, altering prompt templates, relaxing loss at correction markers, and widening the loss window. Evaluation uses four thinking-enabled samples per problem in domain (Avg@4) and 12 samples per problem on AIME 2024, AIME 2025 and HMMT February 2025, with 10,000 problem-cluster bootstrap resamples for paired 95% intervals and Holm correction across the six initial view contrasts.
Why This Matters
Impact on research: The paper reframes how privileged information should be evaluated in self-distillation. It shows that reporting a gain over the base conflates two different effects, and that the value of a reference depends on the student's own responses, the student's training trajectory and the checkpoint at which comparisons are made. The AMPLE-Math suite, its answer-matched design, and the profiling-plus-matched-intervention protocol give future work a way to test claims about references rather than assuming that a fuller solution teaches more.
Real-world applications:
- Post-training open reasoning models where no larger teacher model is available, by making better use of the model's own parameters.
- Selecting and compressing supervision material (gist, key points, polished solutions) for training pipelines where full traces are expensive to store or process.
- Math and STEM tutoring systems, where the question of whether to expose a model to worked solutions versus final answers affects both accuracy and response style.
- Model evaluation practice, since the paper shows that a configuration can look better on one inference mode and worse on another for the same checkpoints.
Industry relevance: Self-distillation is attractive because it avoids the cost of a separate stronger teacher, and these results indicate that much of the benefit arrives without any reference at all, while naive interventions such as relaxing loss at reconsideration markers or extending the loss window may not change behavior or may merely reflect stopping earlier. That has direct implications for how teams design reference data, choose stopping points and interpret benchmark movements.
Future Directions
- Adaptive references: The paper explicitly motivates testing whether reference content that adapts to the student's evolving attempts improves on a fixed representation held constant throughout training.
- Joint reference-and-student design: Treating reference design and student training as one joint problem rather than separable choices, including configurations where reference form is matched to the rollout mode being trained.
- Mechanistic account of cross-mode transfer: The authors interpret gains as improved access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference; identifying and verifying those shared parameters remains open.
- Better protocols against stopping-time confounds: The loss-window result shows checkpoint choice explaining most of an apparent gain, motivating broader adoption of matched-step and development-selected checkpoint comparisons in OPSD research.
Target Audience
Researchers and graduate students working on language model post-training, distillation and self-improvement; practitioners who build or tune reasoning models and need to decide what supervision content to supply; and readers interested in learning with privileged information. Readers who want the high-level lesson can read the abstract, introduction and discussion, while the profiling statistics, matched interventions and appendices are aimed at those who will reproduce or extend the experiments. The authors note their evidence comes from short-run LoRA training on mathematics in two model families, and that AMPLE-Math's release is subject to the licenses of its upstream data and models.
Authors’ abstract
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for an additional reference benefit is modest in Qwen, strongest for a polished solution, whereas complete traces add two percentage points in SmolLM3-3B at step 50. These benefits depend on the student being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families while the problems, references, and evaluation stay fixed. Teacher profiles and matched loss interventions in Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference. The value of a privileged reference is what it adds to this cross-mode transfer, not how much of the solution it reveals.