Research
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
Overview Research area: Large language model training and compression, specifically knowledge distillation across models that use different tokenizers (Cross-Tokenizer On-Policy Distillation), with ev

- arXiv
- 2610.08448
- Published
- 2026-10-06
- Authors
- Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu, Yuewei Zhang
AI summary
Overview
- Research area: Large language model training and compression, specifically knowledge distillation across models that use different tokenizers (Cross-Tokenizer On-Policy Distillation), with evaluation on mathematical reasoning and code generation.
- Technical level: Advanced. The paper assumes familiarity with reverse KL divergence, teacher–student distillation, tokenization schemes, and gradient-based optimization diagnostics.
- Scope: Across three heterogeneous teacher–student model pairs, the paper measures how much strict 1:1 token alignment already covers student-generated text, tests whether adding supervision on non-aligned "mismatch" spans helps or hurts, and shows that a small student-selected top-16 subset of the shared vocabulary can match or beat full shared-vocabulary distillation and the evaluated cross-tokenizer baselines.
What This Paper Is About
When a teacher model and a student model use different tokenizers, the same response is chopped into different tokens and the two models predict over different vocabularies. Prior work treats this as a coverage problem, adding more alignment machinery (span grouping, learned mappings, byte-level interfaces) to supervise more of the response. This paper asks the opposite question: how much useful supervision does strict one-token-to-one-token matching already retain, and does recovering the excluded positions actually improve the student? The authors find that strict alignment already covers most tokens and nearly all probability mass, and that adding supervision on the remaining mismatch spans lowers accuracy.
Key Contributions
-
A coverage-versus-reliability reframing. The paper distinguishes alignment coverage (how much of the response and vocabulary can be compared) from supervision reliability (whether the resulting teacher targets give useful guidance at the states the student actually visits), and argues the field should optimize the latter.
-
Measurement of strict alignment and predictive mass. The authors report strict 1:1 token coverage of 85.57–96.98% of student tokens and 82.82–97.26% of teacher tokens across three runs, despite static vocabulary Jaccard overlap of only 39.49–64.87%, and show the shared vocabulary retains 99.69–99.90% of teacher and 98.99–99.81% of student probability mass at strict positions.
-
A controlled test of mismatch-span supervision. Holding the strict reverse-KL objective fixed, they add a mean squared error loss on span log-probabilities over mismatch groups and sweep the weight λ ∈ {0, 0.25, 0.5, 0.75, 1.0, 1.25, 1.5}. Every positive weight achieves 100% structural supervision coverage yet lowers full-average accuracy.
-
Compact student-selected top-k supervision. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary retains at least 96% of the full-average improvement of full shared-vocabulary OPD and outperforms ULD, Extended ULD, GOLD, and SimCT on all three model pairs.
-
Gradient diagnostics. Checks from the strict-only run at steps 0, 20, 60, and 100 show mismatch-span gradients with weak or negative cosine agreement with strict gradients, a split-half strict reference that stays higher at every checkpoint, and a span-to-strict gradient norm ratio that grows over training.
Main Findings
-
Static vocabulary mismatch does not imply poor alignment. Vocabulary Jaccard overlap ranges from 39.49% to 64.87%, yet strict student-token coverage ranges from 85.57% to 96.98% and teacher-token coverage from 82.82% to 97.26% over complete runs. Granite → Phi has the lowest Jaccard overlap (39.49) but the highest coverage under both tokenizations (96.98–97.12 student, 95.83–97.96 teacher across the reported windows).
-
Coverage is stable across training. Aggregating 51,200 saved responses per run (512 rollouts per iteration over 100 iterations), the range across the steps 1–20, 41–60, and 81–100 windows is at most 1.43 percentage points for strict student-token coverage and 3.75 percentage points for strict teacher-token coverage within each pair.
-
Supervising mismatch spans hurts. All three student–teacher pairs attain their highest full average at λ = 0 (strict supervision only). The 18 positive-weight settings score 0.27–1.20 percentage points below their respective strict baselines, and λ = 0 also gives the highest math and code averages for each pair.
-
Probability mass concentrates on shared vocabulary entries. At strict positions on pre-distillation student rollouts, the full shared vocabulary retains 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average. At k = 16, the student-selected subset retains at least 93.54% of teacher mass and 94.55% of student mass in each pair.
-
A small student-selected subset is nearly sufficient. Selection uses only the student distribution, yet the chosen entries also carry most of the teacher's mass. Expanding from 16 to 128 entries increases retained mass by at most 3.44 percentage points for the teacher and 2.70 for the student.
-
Top-16 matches full shared-vocabulary OPD. Top-16 retains at least 96% of the reported full-average improvement from Base to Strict full in each pair, with no consistent further gain from a larger subset. Strict top-16 achieves Full averages of 32.64 (Qwen → Llama), 42.26 (Granite → Phi), and 47.06 (Granite → Qwen), versus 32.86, 42.49, and 47.35 for Strict full.
-
Strict variants beat the cross-tokenizer baselines. All three strict variants exceed ULD, Extended ULD, GOLD, and SimCT in both math and full averages on every model pair; top-16 leads the strongest alternative in the full average by 0.51–1.05 percentage points. In the Granite → Qwen pair, Strict full and Strict top-128 both reach 54.84 and 55.62 code averages respectively, above GOLD (54.84) and SimCT (53.56).
-
Span gradients conflict with strict gradients. At strict-only checkpoints, mismatch–strict gradient cosines are near zero for Qwen → Llama and Granite → Phi; Granite → Qwen starts with negative agreement and approaches zero later. The split-half strict reference cosine decreases overall during training but remains higher than the mismatch–strict cosine at every measured checkpoint.
-
Relative gradient scale grows. The span-to-strict gradient norm ratio increases across the measured checkpoints in all three pairs, meaning a fixed positive span weight would contribute an increasing gradient norm relative to the strict component.
-
Additional settings favor compact supervision. Appendix D reports distillation from a larger 235B teacher to an 8B student where Strict top-16 achieves the highest Math, Code, and Full averages among the evaluated methods; on ALFWorld, Strict top-16 attains the highest success rate on both seen and unseen splits.
Methodology in Plain English
The authors start from a simple baseline they call Strict full cross-tokenizer OPD: for each response the student generates, they tokenize the decoded text with both the student's and teacher's tokenizers, find positions where one student token and one teacher token cover exactly the same span, restrict both models' next-token distributions to the vocabulary entries the two tokenizers share, renormalize, and apply reverse KL there.
To study coverage, they quantify the fraction of non-special student and teacher tokens that fall into these strict 1:1 groups, and compare it against a static vocabulary Jaccard overlap computed from the two vocabularies.
To study mismatch supervision, they group remaining intervals where one side needs multiple tokens, compute the product of token probabilities for the observed path on each side, and add an MSE loss between the log-probabilities of those span products. They sweep the weight on this extra loss.
To study concentration, they take the student's own softmax probabilities at strict positions, pick the student's top-k shared vocabulary entries, and measure how much probability mass both models place on that subset. They then train with reverse KL restricted to that subset for k = 16 and k = 128 and compare against ULD, Extended ULD, GOLD, and SimCT (all implemented in KDFlow).
To explain the accuracy drop, they compute gradients of the strict and span losses on identical fixed responses and report two statistics: the cosine between them and the ratio of their norms, with a control formed by randomly splitting the strict-loss positions into two halves and measuring the cosine between those halves.
Training uses 20,000 prompts (10,000 from DAPO-Math-17K, 10,000 from CodeForces), 100 on-policy iterations with 512 responses per optimizer batch, AdamW at a learning rate of 1×10⁻⁶ with β₁ = 0.9 and β₂ = 0.98, weight decay 0.0, gradient clipping 1.0, checkpoint interval 10 steps, rollout temperature 1.0 and top-p 0.95, maximum prompt length 2,048 and maximum generation length 8,192. Implementation uses FSDP2 on a single node with 8 NVIDIA H20 GPUs with SGLang as the rollout engine.
Evaluation: mathematics on MATH500, GSM8K, AIME-2024, AIME-2025, AIME-2026, AMC23, and Minerva-Math at mean@32; code on HumanEval, MBPP, and LiveCodeBench at mean@8 (LiveCodeBench release_v6, 1,055 problems). Math and Code are arithmetic means within each domain, and Full is their equally weighted mean.
Why This Matters
Impact on research. The paper challenges the implicit assumption behind a family of cross-tokenizer distillation methods (SimCT, ULD, GOLD, byte-prefix marginalization) that recovering more teacher signal is always better. It provides a concrete counterexample where full structural coverage lowers downstream accuracy, and supplies probability-mass and gradient diagnostics as a way to pre-screen whether added supervision is reliable.
Real-world applications.
- Distilling a large proprietary teacher into a smaller deployable student when the two models come from different families and cannot share a tokenizer.
- Building small domain-specialized models for math or code assistants, where the paper's benchmarks (MATH500, GSM8K, AIME, HumanEval, MBPP, LiveCodeBench) are the standard deployment targets.
- Reducing training compute for distillation pipelines: compact top-16 supervision matches full-vocabulary supervision, which means cheaper KL computations and smaller support sets at every strict position.
- Agentic and tool-use training, since the appendices report a 235B-teacher-to-8B-student setup and ALFWorld success rates.
Industry relevance. The work comes from Alibaba Cloud Computing and studies model pairs spanning Qwen, Llama, Granite, and Phi, which are the model families most commonly mixed in production pipelines. The practical takeaway, that a student-selected top-16 shared-vocabulary support at strict positions is enough, is directly usable by teams that want to distill across these families without adopting heavier alignment machinery.
Future Directions
- Generalize the diagnostic to other objectives. The paper tests one span objective (MSE on span log-probabilities). Whether other mismatch objectives, such as rank-based or byte-prefix marginalization losses, show the same weak gradient agreement is not reported.
- Explain and fix the gradient scaling. The span-to-strict gradient norm ratio grows across the measured checkpoints, but the paper does not report a schedule or reweighting that counteracts this growth. A principled normalization or annealing scheme is an open question.
- Test on more modality and task families. The main study covers mathematical reasoning and code generation with three model pairs; the appendices extend to a 235B teacher and ALFWorld, but wider coverage of languages, long-context tasks, and other tokenizer families is not reported.
- Probe the interaction between support size and teacher quality. Top-16 retains at least 96% of the full-average improvement and top-128 is sometimes slightly better on code, so the conditions under which a larger student-selected support helps remain unresolved.
Target Audience
Researchers and engineers working on knowledge distillation, LLM compression, and cross-model transfer, especially those whose teacher and student models come from different families with different tokenizers. The paper is most useful to readers already comfortable with KL-based distillation objectives and gradient analysis; practitioners building distillation pipelines for math or code models will find the top-16 result and the coverage-versus-reliability framing directly actionable.
Authors’ abstract
On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.