Skip to content
AI.info

Research

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

Overview Research area: Machine learning / large language model training, specifically on-policy distillation (OPD) and the failure mode of runaway response length. Technical level: Intermediate. The

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
arXiv
2609.20511
Published
2026-09-17
Authors
Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao, Xuchao Zhang, Chetan Bansal, Huaxiu Yao, Taylor W. Killian, Weitong Zhang

AI summary

Overview

Research area: Machine learning / large language model training, specifically on-policy distillation (OPD) and the failure mode of runaway response length.

Technical level: Intermediate. The core idea is intuitive (two models can mean the same thing with different stop tokens), but the paper's diagnosis and its four proposed fixes involve token-level probability math and knowledge of how distillation objectives are implemented.

Scope: A single 2026 study that diagnoses termination-token mismatch between base students and post-trained teachers as a cause of length inflation in on-policy distillation, compares four corrections across three model families, and extends the analysis across K2-Horizon post-training stages.

What This Paper Is About

On-policy distillation trains a student model on its own generated text while a stronger teacher scores each token. The authors observe that student responses often grow longer and longer during this training and eventually exhaust the generation budget, producing repetitive or redundant filler even after the correct answer has already appeared. Their core claim is that a large part of this "length inflation" comes from a surprisingly mundane cause: the student and the teacher prefer different end-of-sequence (EOS) tokens for the same semantic decision to stop, so the student's own stopping action gets punished without the teacher's preferred stop token ever being learned.

Key Contributions

  1. Identification of termination-token mismatch as a mechanism. Across Qwen3, Llama, and Gemma, the authors show that student and teacher checkpoints can encode the same semantic stopping decision with different learned token preferences, and that this holds even when both models declare identical EOS sets (Gemma 3 is the clearest case).

  2. A four-way comparison of termination-handling strategies on Qwen3. The authors test (1) shared-set decoding, (2) teacher-side EOS mapping, (3) semantic EOS class, and (4) canonical single-EOS action space, and show that aligning only the decoding stopping set is insufficient, while probability-level corrections all behave similarly and mitigate the inflation.

  3. A stage-wise analysis using K2-Horizon-7B. By distilling the final post-trained checkpoint into students initialized from pretraining, midtraining, and SFT checkpoints, the authors trace how learned termination preferences shift during training and reveal a late-stage length re-inflation that survives termination alignment.

  4. A released implementation including the termination-handling corrections and the evaluation protocols, at https://uncsciml.github.io/opd-eos-website.

Main Findings

  • Student and teacher can disagree on the stop token even with identical declared EOS sets. Table 1 shows the Qwen3 base checkpoint declares <|endoftext|> (151643) while the post-trained checkpoint also recognizes <|im_end|> (151645). Gemma 3 PT and IT both declare <eos> and <end_of_turn> (1, 106), yet the PT student concentrates stopping probability on <eos> while the IT teacher favors <end_of_turn>.

  • The mismatch enters the training signal, not just the decoder. In sampled-token OPD, the coefficient for a student-sampled token e is A_t(e) = log π^E(e|s_t) − log π_θ(e|s_t). When the teacher puts less probability on that particular surface token than the student does, the coefficient is negative, so the student's native EOS is suppressed rather than the teacher's alternative being transferred.

  • The teacher-preferred token is effectively unsamplable. In Qwen3, the student assigns roughly 10^-11 probability to the teacher-preferred <|im_end|> around terminal states, so registering it as a stopping token has little practical effect. Consistent with the mechanism, the student's probability on its native EOS falls from roughly 0.8 early in training to near zero, while response length and clipping increase.

  • Decoding-level alignment alone fails. Fix 1 (registering both <|endoftext|> and <|im_end|> as stopping tokens during rollout, leaving the objective unchanged) follows nearly the same response-length and clipping trajectory as vanilla OPD.

  • Probability-level corrections work. Fixes 2, 3, and 4 reconcile termination at the probability level, keep length and clipping substantially closer to teacher reference lines, and prevent the student's native termination probability from collapsing. The authors use Fix 2 (teacher-side EOS mapping) as the simplest correction for the explicit Qwen3 mismatch and Fix 3 (semantic EOS aggregation) as the default for cross-family experiments.

  • Semantic EOS aggregation generalizes across families. Relative to vanilla OPD it substantially reduces response length and clipping in Qwen3, Llama 3.2, and Gemma 3. Qwen3 and Gemma 3 approach their teacher reference lengths relatively closely, while Llama 3.2 retains a larger residual gap.

  • Recovery dynamics differ by family. With the semantic correction, Qwen3 reaches and maintains a high termination probability quickly, whereas Llama 3.2 and Gemma 3 stay in a low-probability regime longer and recover only later in training. All three show a transient early length increase before decreasing, weak in Qwen3 and more pronounced in Llama 3.2 and Gemma 3.

  • Stage-wise preference shifts in K2-Horizon. The pretrained student primarily favors <|ifm|endoftext|> while the final post-trained teacher favors <|ifm|im_end|>; most of this shift occurs during midtraining, and by the SFT stage the student's preference is already close to the final teacher's.

  • Vanilla OPD can transfer a termination token when it has support. Because the pretrained K2-Horizon student already assigns non-negligible probability to the teacher-preferred token, it can sample that token and receive direct supervision, so termination mass shifts toward the teacher-preferred surface form under vanilla OPD.

  • A late-stage failure remains after alignment. In the Pretrain-to-Final K2-Horizon runs, after an intermediate plateau, response length increases again, termination probabilities approach zero, and almost all rollouts reach the generation budget. This occurs in the corrected run too, where both EOS tokens are already treated as one shared stopping action, so surface token identity cannot explain it. The transition is relatively abrupt under vanilla OPD and more gradual under semantic EOS correction.

  • The correction is largely inert when representations already agree. Applying semantic EOS aggregation to the K2-Horizon midtraining and SFT students, which already place substantial probability on the teacher-preferred token, produces similar length, clipping, and termination dynamics with no consistent downstream degradation.

  • Template and grader choices matter. Cross-template evaluation can reduce observed response length, most clearly when a DAPO-trained model is evaluated with the TTRL template, and performance gains from OPD can persist despite severe length inflation, particularly under TTRL evaluation. DAPO-style evaluation is more sensitive to long redundant continuations.

Methodology in Plain English

The authors start from a configuration that reliably reproduces the problem: distilling a Qwen3-4B teacher (non-thinking mode) into a Qwen3-1.7B-Base student, both sharing a tokenizer and vocabulary, trained on DAPO-Math-17K with single-turn mathematical reasoning prompts. During training they track mean response length and the clipping ratio (the fraction of responses hitting the generation-length budget). For evaluation they use Avg@16 — 16 sampled responses per problem — averaged over AMC23, AIME24, and AIME25.

To diagnose termination behavior directly, they measure the teacher's and student's probabilities at the final position of each student rollout, reporting both individual token probabilities and the total stopping mass q_π(h) = Σ_{e∈E_EOS} π(e|h). They note this total stopping mass is a different quantity from the fraction of rollouts that terminate naturally, and that teacher reference lines in their figures come from teacher generations on the same prompts, not from teacher probabilities at student prefixes.

They then run a controlled ablation of four corrections. Fix 1 changes only which tokens the rollout decoder treats as stopping tokens. Fix 2 sums the teacher's probability over all termination-equivalent tokens onto a canonical student EOS token (with negligible probability retained elsewhere for numerical stability, preserving total mass). Fix 3 collapses all termination-equivalent tokens into one abstract "stop" action for both models and applies the standard sampled-token OPD update to the aggregated probabilities. Fix 4 maps the teacher's EOS mass to a canonical token and removes all other EOS tokens from the student's sampling distribution, renormalizing.

To check generality, they repeat the base-to-post-trained distillation within Llama 3.2 and Gemma 3 — explicitly within each family, not between families. To study how termination preferences evolve, they use K2-Horizon-7B, fix its final post-trained checkpoint as the teacher, and initialize students from the pretraining, midtraining, and SFT checkpoints; this is useful because the termination token IDs and declared stopping set are unchanged across those stages.

Limits on the setup: training allows at most 1,024 prompt tokens and 7,168 response tokens, while independent checkpoint evaluation allows at most 8,192 generated tokens. The default prompt template is TTRL, which requires the final answer in boxed{}; the template comparison adds the DAPO template and a raw-question template that provides only the problem statement. The authors restrict empirical analysis to sampled-token OPD and do not evaluate full-vocabulary OPD because of its substantially higher computational cost. The provided content does not report specific Avg@16 accuracy figures.

Why This Matters

Impact on research. Prior explanations of OPD length inflation — objective-level reverse-KL preferences, teacher supervision degrading with trajectory depth, entropy collapse — all share a precondition: the student must first fail to stop. This paper argues that termination-token mismatch is a lower-level, implementation-adjacent cause that precedes those mechanisms and is removed by a change to the objective alone. The authors frame their results as complementary rather than contradictory: where teacher and student already agree on termination, prior mechanisms remain the relevant explanation; where they disagree, alignment should be established before attributing inflation to an algorithmic cause. It also broadens the phenomenon beyond Qwen3's explicit EOS-set difference, since Gemma 3 shows the same failure with identical declared EOS sets.

Real-world applications (implications drawn from the paper's findings, not claims the paper itself lists):

  • Training smaller or cheaper reasoning models by distilling from stronger post-trained teachers, which is the exact setup the paper studies.
  • Controlling inference cost and latency, since rollouts that exhaust the generation budget consume the full token allowance.
  • Evaluating and comparing distillation recipes fairly, given the paper's finding that template and grader choices modulate measured length and accuracy.
  • Diagnosing degenerate model outputs in deployed math/reasoning assistants, where the paper notes answers are often correct early but followed by repetitive continuation.

Industry relevance. Length inflation directly translates into compute cost per response and into truncated, repetitive output for end users. The paper's central fix — treating functionally equivalent EOS tokens as one shared semantic stopping action during distillation — is a small, contained change to the training objective rather than a new architecture, and the authors release an implementation with the corrections and evaluation protocols. The finding that the correction is largely inert when termination representations already agree lowers the risk of adopting it broadly.

Future Directions

  • Explain the late-stage re-inflation. The K2-Horizon Pretrain-to-Final run re-inflates even after the teacher-preferred termination form has been transferred and even under semantic EOS correction, with almost all rollouts reaching the budget. The authors explicitly leave its origin for future work, noting only that it is qualitatively similar to the length inflation and truncation collapse reported by Luo et al. (2026) without establishing a shared mechanism.

  • Explain cross-family differences in recovery dynamics. Qwen3 recovers termination probability quickly under correction, while Llama 3.2 and Gemma 3 remain in a low-probability regime longer and retain different residual length gaps from their teachers. The curves describe these differences but, in the authors' words, do not by themselves establish their cause.

  • Quantify the trajectory-level continuation bias. Appendix C sketches an argument that local sampled-token OPD omits the future student–teacher mismatch incurred by continuing, removing a nonnegative contribution to the stopping-logit update relative to a trajectory-consistent reverse KL update. The paper presents this as consistent with early length growth but does not resolve it empirically.

  • Reconcile termination mismatch with the other proposed mechanisms. Since the paper frames its findings as complementary to prior work on objective-level preferences, supervision degradation, and entropy collapse, an open question is how these causes interact within a single training run and in what order they should be addressed.

Target Audience

Researchers and engineers who build or study on-policy distillation pipelines for language models, particularly those distilling post-trained teachers into base or smaller students. It is also useful for practitioners debugging runaway generation length, truncation, or repetition in trained reasoning models, and for anyone comparing distillation recipes across prompt templates and graders. Readers need basic familiarity with token-level training objectives and EOS handling to follow the fix comparisons; the diagnostic framing itself is accessible to a broader machine learning audience.

Authors’ abstract

We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify \emph{termination-token mismatch} between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama, and Gemma, the two models can place their stopping probability on different EOS tokens, even when their declared stopping sets are identical. This mismatch can suppress the student's preferred termination action without reliably transferring the teacher-preferred alternative. We show that aligning the decoding stopping set alone is insufficient, while treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigates mismatch-induced length inflation across all three model families. To further understand how termination behavior evolves over training, we study OPD across different K2-Horizon training stages. This stage-wise analysis shows that termination preferences can shift substantially during training, while also revealing a distinct length inflation late in the OPD run that persists beyond termination alignment. Together, these results identify termination mismatch as an important, but not exhaustive, source of OPD length dynamics. We release an implementation incorporating the proposed termination-handling corrections.

Read the original paper