Research
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Overview Research area: Efficient large language model inference, specifically low-bit quantization of the key-value (KV) cache using data-adaptive linear transforms. Technical level: Advanced. The pa

- arXiv
- 2609.38121
- Published
- 2026-09-29
- Authors
- Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler, Dan Alistarh
AI summary
Overview
- Research area: Efficient large language model inference, specifically low-bit quantization of the key-value (KV) cache using data-adaptive linear transforms.
- Technical level: Advanced. The paper is written clearly, but the core construction (WUSH), the near-optimality theorem, and the grouped-query attention algebra require comfort with linear algebra, second-order statistics, and quantization error analysis.
- Scope: The paper adapts an existing closed-form, product-aware transform (WUSH) to KV-cache quantization, proves a near-optimality guarantee for one integer quantizer, and evaluates the resulting method (WUSH-KV) on attention reconstruction error, WikiText-2 perplexity, and four downstream reasoning benchmarks across three Qwen3 models.
What This Paper Is About
During autoregressive generation, every attention layer stores the keys and values of previous tokens so they do not have to be recomputed, and this KV cache grows linearly with context length and batch size. Compressing the cache to very low bitwidths (2–4 bits) saves memory and bandwidth, but naive quantization introduces large errors into attention scores and outputs. The paper's goal is to choose the change of basis applied to cached keys and values so that quantization error hurts the model as little as possible, by building the transform from the second-order statistics of both factors of each matrix product in which the cache participates.
Key Contributions
- A KV-cache adaptation of the WUSH transform (WUSH-KV). For each KV head, the method builds one transform for keys and one for values from calibration-time statistics. The key transform accounts for the query heads that read the cached keys; the value transform accounts for the output-projection blocks that read the cached values. The value transform is folded into the model weights offline, so it adds no online cost.
- A deliberate transform placement. The key transform is applied after headwise normalization (when present) and after RoPE, with a query-side compensation transform. The paper argues that fully folding the key transform into the projection weights would require commuting with RoPE and learned RMS normalization, which leaves only trivial paired sign-flips; placing an unrestricted transform before RoPE instead would make the compensation position-dependent.
- A near-optimality result for the QuEST integer quantizer. Writing the layer-output loss as a quadratic form in the reconstruction perturbation, the paper defines a transform as sensitivity-balanced when
T^{-⊤} H T^{-1}has equal diagonal entries. Under a Gaussian-tail and clipping-alignment condition, Theorem 2 states that every sensitivity-balanced invertibleTsatisfiesE[ℓ(T★)] ≤ (1 + O(α_b^{-2})) E[ℓ(T)]as bitwidthb → ∞, whereT★is the undamped WUSH transform (γ = 0). - A cache-management scheme plus end-to-end integration. A fixed full-precision sink window of the first
S_sinkpositions and a rolling recent window of the newestS_keeppositions are combined with batched flushing of newly eligible entries in multiples ofS_flush. The method is integrated into SGLang using the OSCAR-style percentile-clipped affine quantizer to enable a direct comparison with concurrent work.
Main Findings
- Lower attention-stage reconstruction error. In a controlled ablation on all 36 attention modules of Qwen3-8B, calibrated on 128 FineWeb-Edu sequences and evaluated on 32 disjoint sequences of length 1024 with a prefill-like setting and the QuEST quantizer held fixed, the geometric-mean module-output errors at 2-bit with both keys and values quantized are 0.208 for WUSH and 0.195 for WUSH-A, compared with 0.325 for normalized Hadamard and 0.309 for OSCAR. The largest layerwise reductions occur in the first few attention modules.
- WUSH-A gives only a modest gain over WUSH. The attention-aware Hessian variant (WUSH-A) requires only forward-pass quantities (no backpropagation or automatic differentiation), but its improvement over the simpler WUSH Hessians is small, so the paper uses the simpler variant for the main method.
- Lowest perplexity among the tested transforms at every bitwidth. On Qwen3-8B WikiText-2 with the full-precision baseline at 9.72: WUSH gives 9.73 at 4-bit, 9.82 at 3-bit, and 10.51 at 2-bit. OSCAR gives 9.77 / 10.19 / 13.74; random orthogonal gives 9.77 / 10.18 / 14.22; identity gives 11.36 / 15.29 / 38.08; normalized Hadamard gives 9.77 / 10.06 / 258.29.
- A diagnosed failure mode for the Hadamard transform. The extreme 2-bit perplexity for H is attributed to spikes in the first attention module's key vectors that are concentrated in essentially one coordinate before the transform; the Hadamard transform spreads these spikes into nearly equal-magnitude coefficients close to the midpoints between adjacent 2-bit reconstruction levels, producing dense quantization error.
- Competitive-to-better downstream accuracy at 2-bit. The abstract and introduction report that at 2-bit WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks, scoring higher than OSCAR on all four Qwen3-8B tasks and being competitive at 4B and 32B. The per-task scores, standard deviations, generated-token counts, and budget-exhaustion rates for the 2-bit downstream table are not present in the provided paper content, so specific numbers are not reported here.
- Practical calibration cost. One-time Qwen3-8B calibration takes about 12 minutes and peaks at 39 GiB of GPU memory on a single NVIDIA L40S GPU. The resulting transforms are reused across all cache bitwidths.
- Low storage and online overhead. The transforms are
n_layer · n_kvkey transforms of sized × d, fixed once at calibration and shared by every position; the value-side transform adds no online cost because both halves are folded into weights. The key side costs oned × dproduct per new key and per query. The full-precision windows typically cover less than 1% of the maximum sequence length in the downstream tasks, so they increase the effective cache bitwidth only slightly.
Methodology in Plain English
The starting point is a simple observation: inserting an invertible matrix T between the two factors of a matrix product does not change the product, so you are free to quantize T·X instead of X and undo the transform afterwards. The question is which T to pick.
WUSH answers this in closed form. It takes two ingredients: the Gram matrix M = X Xᵀ of the tensor being quantized (how the cache entries themselves co-vary), and the Hessian H of the output loss with respect to the quantization perturbation (how sensitive the downstream computation is to error in each direction). From these it builds T = c · 𝓗 Λ^{-1/4} Uᵀ Lᵀ, using a Cholesky factorization of the damped Hessian and a symmetric eigendecomposition of the damped Gram matrix, with a damping ratio γ and a scalar c that roughly preserves the Frobenius norm. The construction is invariant to independent positive rescaling of M and H, so normalization never has to be tracked.
For the KV cache, the relevant "downstream" partners are identified directly: a cached key is read only through the query-key product, and a cached value only through the output-projection product. So the key Hessian is 2 Σ_g Q_{h,g} Q_{h,g}ᵀ over the query heads of the group, and the value Hessian is 2 Σ_g W_{O(h,g)} W_{O(h,g)}ᵀ. Calibration runs an unquantized forward pass, accumulates these statistics across sequences (summing rather than storing all activations), and produces one key and one value transform per attention module and KV head with γ = 10^{-2}.
Evaluation proceeds in three layers. First, a controlled ablation isolates the transform by fixing the quantizer (QuEST, per-token, with clipping from the group RMS) and measuring relative squared Frobenius reconstruction error, quantizing keys and values separately and jointly, without any full-precision cache windows. Second, WikiText-2 perplexity on Qwen3-8B with YaRN factor 4, sequences of length 2048 processed in 16-token chunks. Third, downstream reasoning accuracy, where the method is integrated into SGLang with the OSCAR-style percentile-clipped affine quantizer (clipping at a κ-quantile, affine scale and zero point from the clipped range), using κ_K = 0.96 everywhere and κ_V = 0.92 for Qwen3-4B-Thinking-2507 and Qwen3-8B, raised to κ_V = 0.96 for Qwen3-32B. The paper notes these clipping ratios are reused from OSCAR without retuning for WUSH-transformed distributions, which may leave headroom.
Why This Matters
The KV cache is one of the dominant costs of long-context and large-batch LLM serving: it grows linearly with sequence length and batch size, and it is reread at every decoding step. Making 2-bit cache compression actually usable — rather than degrading the model — directly translates into longer contexts, larger batches, or fewer GPUs for the same workload. The paper's distinctive move is to treat the cached key and value as one factor of a specific matrix product and to build the transform from both factors' second-order statistics, rather than relying on generic rotations such as Hadamard transforms or restricting the transform to be orthogonal.
Real-world applications:
- Long-context inference services (document analysis, codebase-wide reasoning, long chat histories) where the cache would otherwise dominate memory.
- High-throughput serving with large batches, where cache bandwidth rather than model weights limits tokens per second.
- On-device or edge deployment, where the KV cache competes with the weights for a fixed memory budget.
- Reasoning and agentic workloads with long chains of thought, evaluated here on AIME 2025, MATH-500, GPQA Diamond, and LiveCodeBench v6.
Industry relevance: the method is implemented in SGLang, a production serving system compatible with paged KV caches, and shares its quantizer with concurrent work (OSCAR) that is also implemented in SGLang, so the comparison is apples-to-apples within the same pipeline. The offline calibration is cheap (about 12 minutes, 39 GiB on a single L40S for Qwen3-8B) and reusable across bitwidths, and the value-side transform is free at inference time, which keeps the online overhead confined to one d × d product per new key and per query.
Future Directions
- Retuning the clipping ratios for WUSH. The paper explicitly reuses OSCAR's tuned
κvalues without retuning them for WUSH-transformed K/V distributions and states this may be suboptimal, leaving potential headroom from WUSH-specific clipping optimization. - Closing the gap between theory and practice. The near-optimality theorem covers the ideal undamped transform (
γ = 0) with the QuEST quantizer, while experiments use damping and, for end-to-end results, a percentile-clipped affine quantizer that does not readily admit the same proof. - Making the attention-aware variant worthwhile. WUSH-A needs only forward-pass quantities but improves on WUSH only modestly here; the paper frames it as a practical alternative for future models or settings where attention-aware sensitivity provides a larger gain.
- Alternative key-transform placements. The paper rejects fully folding the key transform (only trivial sign-flips survive the commutation constraints) and rejects pre-RoPE transforms (position-dependent compensation complicates the attention kernel); whether there is a placement that keeps both full
d × dfreedom and a simple cache representation remains open. - Reconciling local squared error with end-to-end quality. The OSCAR-style quantizer yields better end-to-end quality than symmetric QuEST despite slightly higher average local squared error, a tradeoff the paper observes but does not fully resolve.
Target Audience
Researchers and engineers working on LLM inference efficiency, KV-cache compression, and post-training quantization, particularly those implementing quantization inside serving systems such as SGLang or vLLM. The paper also suits readers interested in the theory of transform-based quantization, since the near-optimality argument for sensitivity-balanced transforms is a self-contained contribution. Readers need a working knowledge of transformer attention, grouped-query attention, RoPE, and basic linear algebra to follow the derivations; the empirical sections are accessible without the theory.
Authors’ abstract
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.