Research
Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
Overview Research area: large language model post-training for mathematical reasoning, specifically on-policy self-distillation (OPSD) and teacher-student supervision design. Technical level: Advanced

- arXiv
- 2609.39687
- Published
- 2026-09-30
- Authors
- Xincheng Wei, Yifan Ding, Yoshua Li, Yuquan Lu, Ziheng Li, Yi Lu, Dongsheng Ma, Rongxiang Weng, Xunliang Cai
AI summary
Overview
Research area: large language model post-training for mathematical reasoning, specifically on-policy self-distillation (OPSD) and teacher-student supervision design. Technical level: Advanced — the paper assumes familiarity with on-policy distillation, privileged-teacher setups, forward-KL objectives, and model routing. One-sentence scope: The paper proposes Neighborhood OPSD (N-OPSD), which builds a pool of perturbed frozen teacher experts and routes among them to give a student model richer supervision at states the student actually visits.
What This Paper Is About
On-policy self-distillation trains a reasoning model with a "privileged" teacher that can see a reference solution while supervising prefixes sampled from the student. Standard OPSD uses a single fixed teacher parameter setting at every step. This paper observes that slightly perturbed versions of that teacher produce different, complementary corrections that still align with the same reference solution, and it aims to collect and route among those corrections so the student receives better-targeted supervision.
Key Contributions
- An empirical observation that local parameter perturbations of the privileged teacher yield complementary, reference-aligned corrections, and that a pool of such experts covers more useful reference positions than the unperturbed teacher alone.
- An offline greedy selection procedure that builds a compact pool of frozen experts by rewarding filtered reference-token gains over the pool's current best at each position.
- An online routing scheme that separates the "anchor" token direction from its level of support: MaxPeak picks the anchor token, and quantile selection chooses among experts whose top token matches that anchor.
- A training pipeline in which the student learns from the chosen expert's full next-token distribution via the clipped forward-KL objective inherited from OPSD, evaluated across three model sizes and three math benchmarks.
Main Findings
- Complementary supervision from perturbations: Local parameter perturbations reveal distinct reference-aligned corrections within the same reference context, and different experts supply these corrections at different reference positions.
- Pool coverage: The expert pool collectively covers more of these reference positions than the unperturbed privileged teacher does.
- Peak is not the best target: The expert with the highest peak need not provide the best training target, which motivates decoupling anchor selection from support-level selection in routing.
- Benchmark gains: Averaged over three independent runs per method, N-OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively, on AIME 2024, AIME 2025, and HMMT February 2025.
- Pool generalization: Student-prefix continuations support using the pool beyond the reference trajectories used during selection.
- Ablation support: Matched ablations support filtered reference-token gains as a selection criterion, and accounting for overlap within the pool plus routing by state further improve student accuracy.
- Deployment cost: Inference uses only the distilled student, so the expert pool is not needed at test time.
Methodology in Plain English
The starting point is a teacher model that is allowed to read a reference solution while it supervises the student on prefixes the student itself generated. Instead of keeping that teacher fixed, the authors create several slightly perturbed copies of it. Each copy makes somewhat different predictions at different points along the reference solution, so together they cover more ground than any single copy.
To keep this manageable, they precompute a small, fixed pool of these expert copies offline. The pool is grown greedily: a candidate is kept when it improves on what the pool already achieves at a given reference position, using a filtered measure of reference-token gains as the score.
At training time, the student visits states, and the system must decide which expert in the pool to learn from at each one. The authors split this into two decisions. First, MaxPeak picks the anchor token — the direction to move toward. Second, quantile selection picks among the experts whose own top token agrees with that anchor, which keeps the direction consistent while choosing how strong the supervision signal is. The student then matches that selected expert's full next-token distribution rather than just its single top token, using the same clipped forward-KL objective that OPSD already uses.
Evaluation is on AIME 2024, AIME 2025, and HMMT February 2025, with three independent runs per method to reduce noise in the comparison. Beyond the headline numbers, the abstract reports that the pool remains useful on continuations of student prefixes and that ablations back the selection criterion, but it does not give the detailed scores or the sizes of the pool, perturbations, or training sets.
Why This Matters
The work reframes teacher quality in self-distillation as a question of diversity and routing rather than a single fixed checkpoint. If a family of nearby teacher settings collectively supplies supervision that no single setting provides, then how supervision is selected becomes a first-class design choice, not an afterthought. The two-stage routing idea also makes a distinction — direction versus confidence — that may transfer to other settings where multiple teachers or reward signals must be combined.
Real-world applications:
- Training smaller, cheaper math-tutoring or homework-help models that inherit reasoning behavior from stronger reference-guided teachers.
- Building verifier-style or step-checking systems for education and technical documentation, where reference solutions are available during training.
- Compressing large reasoning models into deployable sizes for on-device or cost-sensitive assistants, since inference uses only the student.
- Improving any pipeline where a model must be trained against reference solutions, such as code generation with test cases or structured problem solving.
Industry relevance: the method targets the common industrial situation of having strong reference data but wanting a smaller served model. Producing better supervision from an existing teacher setup, without adding anything to the inference path, is directly aligned with cost and latency constraints in deployed reasoning systems.
Future Directions
- Determining how to generate and size the perturbed expert pool in a principled way, rather than relying on the greedy offline procedure described here.
- Testing whether direction-versus-support routing generalizes to other domains beyond mathematical reasoning, and to settings with multiple heterogeneous teachers rather than perturbed copies of one.
- Investigating how pool overlap should be handled in general, since the abstract reports that accounting for overlap improves accuracy but gives no detail.
- Establishing whether the approach scales past the model sizes studied, and how it compares against other forms of multi-teacher or ensemble supervision under matched compute.
Target Audience
Researchers and engineers working on LLM post-training, distillation, and reinforcement-learning-style fine-tuning, particularly those focused on reasoning models. It will also interest practitioners who have reference solutions available at training time and want a smaller student model to benefit from them without paying extra cost at inference. Readers without a background in on-policy distillation, privileged-teacher training, or KL-based objectives will find the routing and selection details demanding.
Authors’ abstract
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.