Skip to content
AI.info

Research

Softmax Reparameterization for Output-Head Quantization

Overview Research area: Model compression and post-training quantization (PTQ) for large language models, focused specifically on the output head (the vocabulary projection that turns a final hidden s

Softmax Reparameterization for Output-Head Quantization
arXiv
2609.31291
Published
2026-09-25
Authors
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King

AI summary

Overview

  • Research area: Model compression and post-training quantization (PTQ) for large language models, focused specifically on the output head (the vocabulary projection that turns a final hidden state into next-token logits).
  • Technical level: Intermediate. The core idea is conceptually simple (softmax only cares about relative logits), but the evaluation involves quantization formats (RTN, AW-MSE, GPTQ, W4/W2, G128 groups), KL divergence, and Fisher-matrix residual analysis.
  • Scope: The paper introduces "softmax reparameterization," a post-training scalar search that picks a functionally equivalent output head before quantization, and evaluates it across seven model heads, three quantizers, plus a packed W4 deployment measurement on Phi-4-mini.

What This Paper Is About

Small language models pair modest decoders with very large vocabularies, which makes the dense output projection a major per-token memory cost: at BF16 it reads about 2Vd bytes per decoding step, versus 0.5Vd in W4. Because softmax depends only on relative logits, many mathematically equivalent output heads exist, and each one produces a different quantization residual. The paper's goal is to exploit that equivalence class: before quantizing, search over shifts that provably leave full-precision predictions unchanged, and pick the shift whose quantized head best matches the source model's predictive distribution.

Key Contributions

  1. Softmax reparameterization. A post-training scalar search over exactly equivalent output-head parameterizations, selected by post-quantization predictive fidelity rather than by weight error. A rank-one correction extends the construction to nonlinear logit paths (such as Gemma 4's tanh soft-capping).
  2. Low-bit recovery across models and quantizers. Across RTN, AW-MSE, and full-Hessian GPTQ at W4, gains concentrate on heads that baseline quantization distorts substantially; at W2, benefits extend across nearly the full model–quantizer matrix. A separate untouched holdout reproduces the improvements on Phi and BLOOM.
  3. Deployment and mechanism. The shift folds into the packed W4 head with no added inference operation for shift-compatible heads, and preserves the latency benefit of head compression. Residual analysis shows that fidelity can improve even when total logit-error energy grows, because probability-weighted relative-logit error falls.

Main Findings

  • W4 gains concentrate where baseline error is large. Test KL falls by 93% on XGLM under RTN (2.13 to 0.143) and by 73–77% on Phi, BLOOM, and BLOOMZ under AW-MSE. Gemma 3/4 and Qwen3.5 already show low W4 error and change little.
  • Full-Hessian GPTQ also improves, from a lower base. On Phi, GPTQ test KL at W4 goes from 0.158 to 0.059; on BLOOM-1.7B from 0.037 to 0.028; on BLOOMZ-1.7B from 0.042 to 0.032.
  • W2 behaves as a compression stress test. Distortion rises across every evaluated head and the benefit broadens, but substantial KL reductions do not restore source-level fidelity in several W2 configurations. On Phi, RTN W2 KL goes from 98.2 to 74.2, AW-MSE from 41.6 to 16.2, and GPTQ from 3.08 to 1.16.
  • Selected shifts beat fixed mean-centering where it matters. On Phi, the validation-selected coefficient is t = 4, lowering RTN test KL from about 0.81 at mean-centering (t = 1) to 0.35, a further 57% reduction. On XGLM, all three quantizers select mean-centering itself (t* = 1). Because the preferred representative depends on the base quantizer, the coefficient is selected separately per quantizer; transferring an RTN-selected coefficient to AW-MSE can reduce fidelity.
  • Mean-centering minimizes weight norm, not quantized KL. Equation 5 shows t = 1 minimizes the full-precision Frobenius norm along this path, but changing t changes groupwise ranges, clipping, and rounding, and therefore changes the quantization residual.
  • Gains survive stronger quantizer controls. On Phi, increasing GPTQ calibration from 1,024 to 65,536 states lowers unshifted W4 test KL from 0.160 to 0.089, and reparameterization further reduces it to 0.035 (61%). Shifting also helps after exact per-channel scaling and with affine RTN using integer zero points; paired article-bootstrap 95% intervals exclude zero for each gain.
  • An independent holdout confirms the result. On 26 previously unused WikiText articles, AW-MSE KL falls by approximately 74% on both Phi and BLOOM, with paired bootstrap 95% intervals excluding zero for all four reductions.
  • Frozen coefficients transfer across domains. WikiText-selected coefficients, applied without retuning, outperform fixed mean-centering in all 18 comparisons on C4 and OpenWebMath where the coefficient differs from 1, and tie the remaining six XGLM cases where it equals 1. A BLOOM stability study with 1,024 validation documents yields identical selected coefficients across ten seeds for each C4/OpenWebMath–quantizer pair, all different from 1.
  • Packed deployment preserves quality gains and speedups. On 16,352 matched WikiText tokens with vLLM and the Marlin INT4 kernel, the shift lowers AW-MSE perplexity from 31.65 to 15.40 and KL to BF16 from 0.97 to 0.28. GPTQ perplexity falls from 13.67 to 12.23 and KL from 0.15 to 0.06, leaving roughly 5% perplexity excess over BF16's 11.65.
  • Latency benefit is retained. Shifted AW-MSE on Phi-4-mini has 10.8% lower batch-one latency and 9.4% lower batch-16 latency than BF16 (A10G, 64-token greedy generation, prefix caching disabled). Min–max and AW-MSE heads show essentially unchanged latency after shifting. GPTQ latency was not measured.
  • The memory cost is deployment-level, not method-level. The tied Phi deployment retains its BF16 input embedding and adds a separate packed output head, raising resident weights from 7.17 to 7.47 GiB — a cost that applies equally to shifted and unshifted W4 heads.
  • Fidelity can improve while reconstruction error grows. On Phi-4-mini over 8,176 test states, moving from t = 0 to t = 4 increases squared logit error 2.8-fold under RTN and 2.1-fold under AW-MSE, while Fisher cost falls by 72–73%.
  • The improvement lands on likely outputs. Under AW-MSE, the high- and middle-probability bins account for 71.8% and 27.9% of the Fisher-cost reduction. Error more than doubles in the lowest-probability bin, which contributes under 0.5% of Fisher cost at either shift. RTN shows nearly the same split: 73.2% and 26.5%.
  • The quadratic approximation tracks measured KL. The Fisher quadratic differs from KL by at most 12.4% across these comparisons (1.4% for RTN and 6.2% for AW-MSE at t = 4). For GPTQ with 65,536 calibration states, the selected t = 4 raises logit-error energy 2.66-fold over t = 0 but lowers Fisher cost and KL by 61%, and the quadratic matches KL within 0.7% on 8,176 recaptured test states.

Methodology in Plain English

The starting observation is that softmax is invariant to adding the same constant to every logit. So if you subtract the same vector from every row of the output weight matrix, the model's predictions are exactly unchanged. This gives a family of equivalent heads parameterized by one scalar, t, along the direction of the vocabulary-row mean: W_t = W − t·1·μᵀ, where μ is the mean of the vocabulary rows.

The authors then quantize several members of this family, including the original head (t = 0) and ordinary mean-centering (t = 1), and evaluate each quantized head on held-out text by KL divergence to the original model's next-token distribution. The t with the lowest validation KL is selected. Because 0 is in the candidate grid, the selected shift can never be worse than the unshifted head on validation KL under the quantizer used for selection — a guarantee the authors explicitly note does not extend to unseen data or a different quantizer.

Experimental protocol: frozen BF16 decoders; only the output head is quantized. Data comes from WikiText — 128 articles with eight states each (1,024 states) for quantizer fitting, and 16 articles each for validation and test, with prefixes of at most 512 tokens. The coefficient grid is 14 points over [-2, 8]: (-2, -1, -0.5, 0, 0.5, 1, 1.5, 2, 2.5, 3, 4, 5, 6, 8), with finer spacing near the two baselines. Each base quantizer selects its own coefficient. Three quantizers are used — RTN, activation-weighted MSE (AW-MSE), and full-Hessian GPTQ — at G128 group size. Full-precision equivalence is verified for every head: linear-softmax heads reproduce the source to KL ≲ 10⁻⁹, while Gemma 4 uses the rank-one soft-cap correction (shifted BF16 KL ≈ 7 × 10⁻¹⁰, source PPL 66.41). The four-head RTN/AW-MSE search takes about five minutes on one A100; Phi's 14-point GPTQ sweep takes about twelve minutes.

For nonlinear logit paths, a rank-one correction restores the terms removed by the shift before the elementwise nonlinearity, so equivalence is preserved at the cost of one extra operation. For shift-compatible (linear-softmax) heads, the shift is folded into the weights and adds no inference operation, and tied models keep their source input embedding while quantizing a separate output copy.

Selection uses KL rather than perplexity, because the authors note perplexity can improve through a change in confidence even when predictions move further from the source.

Why This Matters

Impact on research. The paper reframes output-head quantization: an output head's apparent precision requirement can depend on its parameterization, not only on the function it represents. It sits alongside function-preserving quantization work such as SmoothQuant, AWQ, QuaRot, and SpinQuant, but targets the additive symmetry specific to the output head, in post-training, with selection driven by measured predictive fidelity rather than reconstruction error. The residual analysis — showing that better fidelity can coexist with larger total logit error — offers a diagnostic that reconstrution-based objectives alone cannot express.

Real-world applications:

  • On-device and edge inference for small language models, where single-token decoding is memory-bound and the output projection dominates per-token weight traffic.
  • Multilingual deployments, where vocabularies are large — the paper notes modern SLMs deploy vocabularies far larger than historical 32K-scale designs, up to the 262.2K reported for Gemma and 256.0 for XGLM.
  • Serving stacks that deliberately keep the head at higher precision. The paper observes that practical PTQ recipes often retain the output head at higher precision, leaving a substantial per-token weight read even when the decoder is aggressively compressed.
  • Tied-embedding models, where the head shares its source matrix with the input embedding — the paper's approach retains the BF16 input embedding and quantizes a separate output copy.

Industry relevance. The method is post-training, requires no decoder retraining, and is compatible with existing packed low-bit kernels (the evaluations use Marlin INT4 through vLLM). It preserves the latency benefit of head compression while recovering fidelity, and adds no inference operation for linear-softmax heads — properties that fit deployment pipelines rather than research-only settings. The reported 10.8% batch-one and 9.4% batch-16 latency reductions against BF16 on Phi-4-mini are measured, though GPTQ latency was not measured.

Future Directions

  • Wider and stronger coefficient search. The paper restricts the shift to the one-dimensional vocabulary-row-mean direction and says it does not claim this path contains the optimal representative. Appendix J explores a higher-capacity group-wise extension with additional gains under W4 RTN and AW-MSE in a nine-model study and a separate SmolLM3 evaluation with matched search budgets.
  • Whether a preferred coefficient generalizes. Larger C4 and OpenWebMath validation sets can select different coefficients from WikiText, so the preferred coefficient is not universal. It also depends on the base quantizer, and transferring an RTN-selected coefficient to AW-MSE can reduce fidelity.
  • Extending guarantees beyond the selection setting. The validation-KL guarantee holds only for the quantizer used for selection; it does not extend to unseen data or a different quantizer.
  • Calibration and refit sensitivity. The reported ablations quantify evaluation-sample uncertainty and coefficient-selection stability, but the paper states they do not measure sensitivity to alternative calibration samples or refitted quantizers.

Target Audience

Machine learning engineers and applied researchers working on inference optimization and quantization for language models — particularly those deploying small language models with large vocabularies, using packed low-bit kernels such as Marlin or vLLM, or handling tied-embedding architectures. The paper is also relevant to quantization researchers interested in function-preserving reparameterizations and in why reconstruction error and predictive fidelity can diverge.

Authors’ abstract

Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from $1$ and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.

Read the original paper