Skip to content
AI.info

Research

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Authors: Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye Affiliations: Shanghai

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
arXiv
2609.33463
Published
2026-09-27
Authors
Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye

AI summary

Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features

Authors: Cunchun Li, Haonan He, Yifan Gao, Minglei Li, Jingqi Ye, Qingyu Yang, Peng Ye Affiliations: Shanghai AI Laboratory; University of Science and Technology of China; Fudan University; KTH Royal Institute of Technology; The Chinese University of Hong Kong arXiv: 2609.33463v1 [cs.CL], 27 Sep 2026 | License: CC BY 4.0

Overview

Research area: Natural Language Processing / large language model post-training — specifically supervised fine-tuning (SFT), token-level loss reweighting, and post-hoc adaptation of learned fine-tuning updates.

Technical level: Advanced. The paper builds a policy-gradient-style loss formalism, derives an entropy-advantage lemma, proves a local one-sidedness proposition for weighted cross-entropy, and states a consistency assumption with a convergence theorem.

Scope in one sentence: The paper analyzes the shared nonnegative-coefficient structure of existing token-reweighting methods for SFT and proposes SCALE, a post-hoc, entropy-guided gating method that suppresses, reverses, or extrapolates a frozen SFT delta, evaluated across three backbones on mathematical reasoning, general-capability retention, and code generation.

What This Paper Is About

Standard supervised fine-tuning updates a model most aggressively on the demonstrated tokens it finds least likely, which helps it acquire new behaviors but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. The paper shows that existing token-reweighting fixes (DFT, ProFit, EAFT, and related approaches) all assign nonnegative coefficients to demonstrated tokens, so they can only suppress or amplify the supervised update — they cannot reverse a harmful feature once learned, and larger weights are not the same as extrapolating a fixed SFT direction. The goal is therefore to keep the completed SFT delta frozen as a stable reference frame and instead learn how strongly that delta should act, per token and per module, using predictive entropy as the only training signal.

Key Contributions

  1. A unified policy-loss analysis of token reweighting for SFT. The authors show that representative methods such as DFT, ProFit, and EAFT can be expressed as demonstration-action objectives with nonnegative reference-token coefficients (with standard SFT setting w_t = 1 and data filtering covering the case w_t ∈ {0,1}).

  2. Identification of a structural distinction between training-time reweighting and post-hoc feature correction. Token weights can suppress or amplify supervised gradients but cannot reverse learned harmful features or extrapolate beneficial features within a fixed base-to-SFT coordinate system, because they change the optimization trajectory rather than scaling a completed delta.

  3. SCALE (Selective Control of Adaptation via Local Entropy). An entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded gates λ_(ℓ,t) ∈ [λ_min, λ_max] — with λ_min < 0 and λ_max > 1 — to suppress, reverse, and extrapolate frozen SFT features using predictive entropy alone.

  4. Empirical validation across three model settings. Consistent improvements over standard SFT and strong token-reweighting baselines on mathematical reasoning and code generation, plus ablations confirming the importance of extending the correction range beyond suppression.

Main Findings

  • Mathematical reasoning improves on all three backbones. SCALE achieves Avg@16 averages of 37.84 (Qwen2.5-Math-1.5B), 43.60 (Qwen2.5-Math-7B), and 36.57 (Qwen3-4B-Base), exceeding DFT by 5.03, 6.08, and 1.91 points respectively, and obtaining the best result on every benchmark in the comparison.

  • Per-benchmark math scores for SCALE. Qwen2.5-Math-1.5B: Math500 71.60, Minerva 31.44, OlympiadBench 31.34, AIME24 8.55, AMC23 46.25. Qwen2.5-Math-7B: 76.97, 35.51, 36.12, 12.51, 56.88. Qwen3-4B-Base: 69.85, 25.41, 32.77, 9.99, 44.84.

  • General-capability retention stays competitive. SCALE records the highest retention average on both Qwen2.5-Math backbones (33.34 on 1.5B and 45.40 on 7B) and the second-highest on Qwen3-4B-Base (48.34). Its BBH scores are notably high at 38.67 (1.5B) and 64.52 (7B), while OpenBookQA drops to 30.80 on the 7B model relative to the base model's 34.40.

  • Best code-generation averages for all three models. Across HumanEval, HumanEval+, and MBPP, SCALE reaches 34.72% (Qwen2.5-Math-1.5B), 47.35% (Qwen2.5-Math-7B), and 47.34% (Qwen3-4B-Base), the highest average of all compared methods in each case.

  • A zero-cost initialization. With zero gate weights and biases, the gated model recovers the exact Stage-I model; the experiments use a deterministic 10⁻⁶-scale perturbation and check initialization parity.

  • Gate-range preferences are task- and backbone-dependent. High-scoring mathematical configurations occupy the negative-L region; the two Qwen2.5-Math models also favor upper bounds above one, whereas Qwen3-4B-Base reaches its mathematical peak with a smaller upper bound. For coding, the high-scoring region lies near L = 0 with larger positive upper bounds.

  • Both reversal and extrapolation contribute. At fixed U = 1, negative lower bounds improve math on all three backbones, and larger U adds gains for both Qwen2.5-Math models. Expanding [0,1] to [0,6] improves code pass@1 on all three backbones, while even [0,1] already exceeds LoRA in math.

  • Theoretical support is conditional. Lemma 1 shows that minimizing predictive entropy yields the signed full-action advantage A_t^ent(a) = log π(a|c_t) + H_t; a binary margin analysis shows that when the margin is positive, entropy descent strengthens helpful contributions and weakens harmful ones. Theorem 1 shows that, under Assumption 1 (gate-subspace entropy–task consistency, ‖e_ψ‖ ≤ ρβ‖∇_ψ V_k‖ for 0 ≤ ρ < 1), gate-only entropy descent increases the proxy task value V_k.

  • Nonnegative weights are provably one-sided. Proposition 1 states that a nonnegative token weight's direct logit descent direction is w_t(e_(y_t) − π_t), which raises the demonstrated logit, lowers competing logits, or vanishes — it cannot reverse that CE direction. Proposition 2 shows that if the minimizer of a differentiable F on [0,1]^M has nonzero gradient, a point in (L,U)^M with L < 0 < 1 < U has strictly smaller F.

  • Not reported: wall-clock training time, GPU/compute budget, and the total number of training tokens for either stage are not reported in the provided content.

Methodology in Plain English

The approach splits adaptation into two separate stages.

Stage I — learn the delta the ordinary way. A LoRA adapter with rank r = 16, scaling factor α = 16, and zero dropout is trained on all linear layers with the standard SFT objective for one epoch. Training uses total batch size 128, AdamW, cosine decay, and 10% warmup, with a peak learning rate of 5 × 10⁻⁴ for math or 5 × 10⁻⁵ for code. The result is a frozen task delta.

Stage II — calibrate how strongly that delta acts. The pretrained model, the output head, and the learned LoRA factors are all frozen. Only a tiny scalar gate per targeted module is trained. Each gate is a linear projection of the (stop-gradient) hidden state, squashed through tanh with temperature τ, then mapped into a bounded interval (L, U) where L < 0 < 1 < U. The gated output is the frozen base contribution plus λ times the frozen delta contribution, so λ = 1 preserves the original SFT update, 0 ≤ λ < 1 suppresses it, λ < 0 reverses it, and λ > 1 extrapolates it. Because the delta operators never change, the meaning of λ — and therefore the distinction between suppression, reversal, and extrapolation — stays fixed throughout calibration.

The only training signal is predictive entropy. Stage II minimizes a top-k conditional-entropy objective (k = 100) over response positions, with no cross-entropy, KL, or budget term, and no relearning of the delta direction. The authors show this gradient equals a covariance between logits and the local intervention tangent — signed policy credit transmitted through delta controls. They justify entropy as an unsupervised local selection signal, not a correctness oracle, under an explicit alignment condition between entropy advantage and task value within the gate subspace; they also present a binary margin argument showing that when the margin is positive, entropy descent strengthens helpful delta contributions and weakens harmful ones.

Implementation cost. Each gated module adds only d_m + 1 trainable parameters, so the total Stage-II parameter count is Σ_(m ∈ M)(d_m + 1), a small fraction of the r(d_m + d'_m) parameters per module in the Stage-I LoRA adapter.

Evaluation. Three backbones — Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base — against LoRA, DoRA, ProFit, DFT, EAFT, and TALR. Math fine-tuning follows DFT in using NuminaMath-CoT, with evaluation on Math500, Minerva, OlympiadBench, AIME24, and AMC23, plus retention checks on MMLU-Pro, BBH, and OpenBookQA. Code fine-tuning follows MoE²-LoRA in using the Magicoder-OSS corpus, with evaluation on HumanEval, HumanEval+, and MBPP. The authors sweep all 42 combinations of L ∈ {−6, …, 0} and U ∈ {1, …, 6}, training a separate gate for each interval; the "Ours" rows in the main tables report the best intervals from that sweep.

Why This Matters

Impact on research. The paper reframes SFT correction as a question of how to use an already learned residual rather than only how to learn it. It gives a concrete structural argument for why the whole family of nonnegative token-weighting methods has a ceiling — they can never flip the sign of a demonstrated-token update — and it separates training-time trajectory modification from post-hoc control of a completed update, a distinction the authors position against post-training model-editing work such as Task Arithmetic, TIES-Merging, AdaMerging, and WEMoE, which compose or merge multiple frozen task vectors rather than modulating a single one per token and module.

Real-world applications.

  • Mathematical reasoning assistants: improving Avg@16 math scores on competition-style benchmarks such as AIME24, AMC23, OlympiadBench, Minerva, and Math500 without a full retraining cycle.

  • Code assistants: raising pass@1 across HumanEval, HumanEval+, and MBPP where the same adaptation pipeline is reused.

  • Continual or domain-specialized deployment: recovering from an SFT run that has degraded general capabilities, using a cheap post-hoc calibration stage that leaves the base model and adapter untouched.

  • Low-resource fine-tuning shops: because Stage II adds only d_m + 1 parameters per module and requires no delta relearning, correction becomes feasible for teams that can afford one SFT run but not repeated retraining.

  • Industry relevance. The two-stage design fits existing post-training pipelines: SFT produces the delta, and SCALE calibrates it afterward without invalidating prior work or requiring the base weights to be re-tuned. Off-the-shelf LoRA adapters remain reusable artifacts, and the entropy-only Stage-II objective needs response prefixes but no new labels.

Future Directions

  • Understanding when the alignment condition holds. Assumption 1 requires entropy and task value to agree directionally within the gate subspace. Characterizing which data, tasks, and backbones satisfy this — and what happens when the margin is negative and the conclusion reverses — is left open.

  • Explaining the task split between math and code. The paper observes that math benefits from reversal and code from extrapolation, and suggests this may reflect how Qwen2.5-Math's specialization versus Qwen3-4B-Base's general-purpose training align with the learned delta. That hypothesis is not resolved.

  • Moving beyond a single frozen SFT delta. SCALE modulates one completed delta; extending token- and module-conditioned signed scaling to multiple deltas, or to combining it with merging-style techniques, is an explicit point of contact with prior work the authors leave unexplored.

  • Reducing reliance on the gate-range sweep. Results depend on selecting from 42 endpoint pairs; the selected maxima are reported as oracle sweep results, and whether the best interval can be predicted a priori for a new backbone or task is not established.

  • Beyond bounded gates. The bounded range limits module-level intervention magnitude but is explicitly not an output-divergence guarantee, and the nonlinear optimum is not guaranteed to lie at L or U — leaving the question of optimal intervention parameterization open.

Target Audience

Most valuable for: researchers and engineers working on LLM post-training, SFT loss design, and parameter-efficient adaptation who are already comfortable with policy-gradient formulations and LoRA mechanics. Also relevant for: practitioners building math or code reasoning systems who need a cheap, post-hoc way to correct a fine-tuning run that hurt general capabilities. Less suited for: readers seeking a beginner-level introduction — the derivations (Lemma 1, Propositions 1 and 2, Assumption 1, Theorem 1) and notation table assume prior familiarity with policy-loss surrogates, entropy advantages, and delta-based weight arithmetic.

Authors’ abstract

Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.

Read the original paper