Research
Mitigating the Length-Scaling Tax with Online Distillation
Mitigating the Length-Scaling Tax with Online Distillation Overview Research area: Reinforcement learning post-training for large language models (RLVR), reasoning efficiency, and on-policy distillati

- arXiv
- 2609.38854
- Published
- 2026-09-30
- Authors
- Xu Wan, Wenyue Xu, Shengjie Zhao, Mingyang Sun
AI summary
Mitigating the Length-Scaling Tax with Online DistillationOverview
Research area: Reinforcement learning post-training for large language models (RLVR), reasoning efficiency, and on-policy distillation.
Technical level: Intermediate. The paper is formal — it defines a metric, a mixed RL/distillation objective, and three KL-based variants — but the motivation and experimental story are readable with a basic grasp of policy-gradient RL and KL divergence.
One-sentence scope: The paper defines the "length-scaling tax" (LST) — extra response length that RL post-training adds to already-solved queries without accuracy gains — and proposes Length Self-Distillation (LSD), which routes solved rollout groups to self-distillation and keeps RL for unsolved groups.
What This Paper Is About
RL post-training on verifiable rewards (RLVR) is widely used to improve LLM reasoning, and the spontaneous growth of response length is often read as a sign of better reasoning. The authors show a side effect: as one shared policy is trained on prompts of mixed difficulty, responses to already-solved "easy" queries get longer even though accuracy on those queries does not improve. They call this wasted growth the length-scaling tax (LST), and their goal is to preserve concise behavior where the policy already succeeds without restricting exploration on problems it has not yet solved.
Key Contributions
-
Defining LST. The paper introduces a query-conditional metric that measures excess response length on already-solved prompts relative to an accuracy-qualified reference, computed on a frozen easy-query set to avoid survivorship bias. It establishes the prevalence of LST under standard RLVR, characterizes its behavioral signatures, and identifies key drivers.
-
Proposing LSD. Length Self-Distillation routes solved rollout groups to an on-policy distillation (OPD) objective while retaining the original RLVR objective for unsolved groups. The teacher is a delayed version of the same policy lineage, instantiated as a rolling exponential moving average (EMA) checkpoint — no external teacher or separately prompted concise model is required.
-
Three implementations with a granularity analysis. The paper develops three LSD variants — supervised-gradient forward KL (SG-FKL), supervised-gradient reverse KL (SG-RKL), and policy-gradient reverse KL (PG-RKL) — and analyzes how they constrain the policy at different levels of granularity.
-
Evaluation across task types and ablations. LSD is evaluated on single-turn mathematical reasoning and multi-turn agentic tasks, with ablations on the distillation objective, EMA teacher half-life, and routing threshold.
Main Findings
-
LST emerges under standard RLVR. Training Qwen3-4B-Base on deduplicated DAPO-Math-17K with a 4096-token response budget and evaluating on AMC 2023, AIME 2025, and AIME 2026 (32 responses per query), aggregate accuracy improves while mean response length grows steadily. When easy sets are frozen at anchors b ∈ {0, 25, 50, 75, 100} using a solve-rate threshold τ = 0.875, accuracy stays largely stable while response length keeps increasing.
-
The extra tokens are not valuable. Analysis of 5,824 responses from the 26 queries in the step-100 easy set shows that length growth is accompanied by a substantial increase in reflection words and repeated bigrams — undesirable behavior on easy queries.
-
Hard-data training amplifies the tax. In controlled comparisons, Hard-4k (4k budget, restricted to prompts the base model solves in at most four of eight rollouts) produces the highest LST across all easy sets, with values of 53.9, 58.0, 54.7, 49.7, and 54.7 across the five anchors at step 400, against All-4k baselines of 42.9, 44.6, 43.0, 40.0, and 43.4.
-
Budget timing matters asymmetrically. All-8k (8192-token budget from the start) amplifies LST strongly at step 240 — e.g., 40.9 vs 18.1 on the anchor-25 set, an increase of 22.8. By contrast, All-4k-8k (4k first, then continued at 8k) shows lower LST than continued 4k training at step 400, with a larger Pass@1 improvement (+15.0 pp vs +11.8 pp), even though its overall and hard-query responses remain longer.
-
A zero gradient from easy groups does not freeze their behavior. The paper's analysis shows that when all responses in a group are correct, group-relative advantages collapse toward zero, so hard groups dominate the update; yet because all prompts share parameters, updates from hard prompts still change token probabilities at easy prefixes to first order.
-
LSD keeps hard-query reasoning long. Hard-query average length increases from 2329 tokens under RL to 2511, 2395, and 2393 tokens under SG-FKL, SG-RKL, and PG-RKL — shorter easy responses and longer hard responses at the same time.
-
Routing proxy is validated. Easy prompts selected at step 500 (thresholds τ ∈ {0.8, 0.9, 1.0}) are traced back through earlier checkpoints; over 87% of prompts classified as easy at step-500 are also classified as easy at step-250.
-
Single-turn LST reduction. LST_1000 falls from 19.03% under RL to −3.72% (SG-FKL), −10.92% (SG-RKL), and 1.45% (PG-RKL). Average Pass@1 is 44.20% (SG-FKL), 42.80% (SG-RKL), and 43.56% (PG-RKL) versus 42.90% for RL; Pass@32 is 67.13% for all three LSD variants versus 65.73% for RL. Reference length for LST_1000 is 853.716 tokens, the mean length of RL step 885, the shortest RL checkpoint satisfying the 100% accuracy constraint.
-
Multi-turn agentic LST reduction. On BrowseComp-Plus, LST_50 falls from 31.4% under RL to 13.7% (SG-FKL), 9.2% (SG-RKL), and 16.1% (PG-RKL). Pass@1 is 30.72%, 29.60%, and 31.65%, versus 29.50% for RL. SG-FKL and SG-RKL also reduce average turns (5.24 and 4.93 vs 5.41), while PG-RKL uses more (5.66) — lower LST does not necessarily imply fewer interactions.
-
Baselines behave differently. CRISP reaches 40.56% Pass@1 with LST_1000 of 10.60%. Fixed SG-FKL (frozen π₀ teacher) collapses to 30.44% Pass@1 despite LST_1000 of −10.81%. On multi-turn tasks, RL + Length Penalty drops LST_50 to 1.2% and average turns to 4.66 but lowers Pass@1 to 22.58%.
-
Training stays focused on hard queries. After step 500, the easy route accounts for 12.17%, 10.58%, and 10.94% of tokens under SG-FKL, SG-RKL, and PG-RKL; the hard route retains 87.83–89.42% of the logged token share on average.
-
Objective granularity has measurable consequences. Among the EMA variants, SG-RKL achieves the lowest LST but its stronger easy-query compression comes with lower Pass@1. SG-RKL has lower mean actor entropy (0.0367) than SG-FKL (0.0743) and PG-RKL (0.0542), and the smallest mean absolute teacher–rollout log-probability difference before updates.
-
EMA half-life trade-off. Comparing H ∈ {2, 4, 8} at τ = 1, H = 4 gives the highest average Pass@1 at step 1750 (41.85%), while H = 8 suppresses LST more strongly (LST_0 = 13.01, LST_500 = −6.21) but slows capability acquisition (40.49%). A longer half-life produces a greater parameter lag and higher distillation loss while routing a smaller share of tokens to OPD.
-
Routing threshold trade-off. Lowering τ from 1.0 to 0.85 and 0.75 routes more partially solved groups to distillation and reduces average Pass@1 from 41.85% to 40.91% and 39.31%; the mean distillation token share rises from 10.31% to 19.00%, and the effective OPD coefficient from 0.247 to 0.450. This supports restricting preservation to fully solved rollout groups, which lack a group-relative reward signal.
Methodology in Plain English
Defining the tax. For each evaluation query, the authors sample N responses from the policy at a given checkpoint and record the empirical solve rate and mean response length. They pick an "anchor" checkpoint, select the queries the policy solves reliably at that point (threshold τ, default 1), and then freeze that set. Evaluating later checkpoints on the frozen set gives a clean before/after comparison: any accuracy change and any length change are measured on the same queries. The tax is the percentage by which the current mean length exceeds the shortest mean length achieved by any RL checkpoint that still meets the accuracy criterion.
Characterizing drivers. They run controlled comparisons varying the rollout budget (4096 vs 8192, and 4096 then 8192) and the training distribution (full DAPO-Math-17K mixture vs hard-only prompts), reporting LST at five anchor-defined easy sets plus overall Pass@1 improvement.
The algorithm. At each training step, the current policy generates G responses per prompt in the batch. The empirical solve rate of the group decides the route: groups solved at or above the threshold go to an on-policy distillation loss; unsolved groups stay on the original RLVR objective. The two route losses are combined by weighting each in proportion to the number of response sequences it contains.
The teacher. Rather than a stronger external model or a concise-prompted copy, LSD uses an exponential moving average of the online policy's own weights, θ̄_k ← βθ̄_{k−1} + (1 − β)θ_k, with β set through a half-life H via β = 2^(−1/H). A larger H means the teacher lags further behind and acts as a stronger behavioral anchor. The authors use the student's own rollout solve rate as a cheap routing signal instead of running extra teacher rollouts, and validate that proxy historically.
The three objectives. SG-FKL minimizes the teacher-to-student KL using the teacher's top-K tokens (K = 32) plus stop tokens, with the teacher distribution normalized over that support. SG-RKL minimizes the student-to-teacher KL over the same augmented support. PG-RKL uses the teacher-minus-rollout-policy log-probability difference at each sampled token as a detached advantage in a PPO-style objective. The first two supervise the full retained support; the third updates only sampled actions.
Evaluation. Single-turn experiments post-train Qwen3-4B-Base on deduplicated DAPO-Math-17K, evaluating AMC 2023 and AIME 2025–2026 with 32 responses per query and a 4k budget, against RL, CRISP, and Fixed SG-FKL. Multi-turn experiments follow Wu et al. (2025), post-training Qwen3-8B-Base on the CutTheBill training split with a 20,000-token response budget and at most 48 turns, using a separate Qwen3-8B refinement agent and evaluating on BrowseComp-Plus against RL and RL + Length Penalty. Training routing thresholds are separate from fixed-easy-set evaluation thresholds.
Why This Matters
Impact on research. The paper reframes length growth during RL not as a uniform signature of better reasoning but as an asymmetric cost: useful on hard problems, wasteful on solved ones. Its metric conditions on a frozen, accuracy-qualified query set, which separates this from coarse global averages of response length and from survivorship-biased easy-set redefinition. It also connects a known failure of group-relative objectives — zero advantage when every rollout is correct — to an observable behavioral drift, and shows that a modest algorithmic addition can counter it without an external teacher.
Real-world applications (the paper does not itself list deployments; these follow from the benchmarks and task settings it evaluates):
- Mathematical reasoning assistants, where the paper's single-turn results come from AMC 2023 and AIME 2025–2026 style competition problems.
- Agentic web-search or research assistants, mirroring the BrowseComp-Plus multi-turn setting with tool observations and a refinement agent.
- Systems where per-query token cost is the binding constraint — short easy queries stay cheap while hard queries keep their full deliberation budget.
- Any deployment serving mixed-difficulty traffic, where a single trained policy handles both routine and difficult requests.
Industry relevance. The reported token allocation — over 87% of training tokens still going to hard prompts after step 500 — indicates the method preserves most of the compute spent on capability acquisition while cutting waste on solved queries. That is directly relevant to serving cost, latency, and interaction-turn counts, and the EMA teacher requires no additional model to host.
Future Directions
-
Scaling to larger models. The conclusion states that future work will evaluate LSD on larger models; all experiments here use Qwen3-4B-Base and Qwen3-8B-Base.
-
Denser learning signals from sampled data. The authors propose extracting richer supervision from sampled rollouts for more effective training.
-
Understanding which training signal drives ablation effects. The paper notes that the coupled changes in sequence share, token share, and effective OPD coefficient across thresholds "do not isolate which training signal causes the performance difference."
-
Reconciling LP-style length penalties with accuracy. RL + Length Penalty achieves the lowest LST_50 (1.2%) and fewest turns (4.66) but drops Pass@1 to 22.58% from RL's 29.50%; finding a penalty schedule that avoids this cost remains open.
-
Detecting waste without an accuracy threshold. The paper notes that for settings reported with τ < 1, LST measures excess length under an accuracy threshold and "does not by itself establish waste at identical accuracy."
Target Audience
Researchers and engineers working on RL post-training, reasoning efficiency, and inference-cost reduction for LLMs. It is most useful to readers already comfortable with policy-gradient methods, group-relative advantages, and KL-based distillation, and to practitioners who need to serve mixed-difficulty query traffic under fixed token budgets. Readers looking for a purely empirical efficiency survey will find the metric definitions and the negative-LST reporting conventions (explained in the appendices) require close attention.
Authors’ abstract
Length scaling during reinforcement-learning (RL) post-training is often viewed as a sign of improved reasoning ability, especially on difficult problems, but may also make responses to already-solved problems unnecessarily verbose. We quantify this side effect as the length-scaling tax (LST): excess response length on already-solved queries without a commensurate accuracy gain. To mitigate LST, we propose Length Self-Distillation (LSD), which routes solved prompts to on-policy distillation and retains the original RL objective for unsolved prompts. LSD uses an exponential moving average of the online policy as its teacher, requiring no external model. We find that LSD achieves comparable or better performance than RL across multiple variants, while substantially curbing response-length growth on easy queries. LSD reduces LST from 19.0% to -3.7% on single-turn reasoning and from 31.4% to 13.7% on multi-turn agentic tasks, demonstrating that LSD effectively preserves concise response patterns on easy queries while supporting efficient exploration on difficult queries during RL post-training.