Skip to content
AI.info

Research

Disaggregated Quantization: Specializing LLM Prefill and Decode

Disaggregated Quantization: Specializing LLM Prefill and Decode Overview Research area: Efficient LLM inference — quantization, low-precision compute kernels, and disaggregated (prefill/decode-split)

Disaggregated Quantization: Specializing LLM Prefill and Decode
arXiv
2609.26333
Published
2026-09-22
Authors
Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi, Tijmen Blankevoort, Dan Alistarh

AI summary

Disaggregated Quantization: Specializing LLM Prefill and Decode

Overview

Research area: Efficient LLM inference — quantization, low-precision compute kernels, and disaggregated (prefill/decode-split) serving systems. The work sits at the intersection of model compression, quantization-aware training/distillation, and inference-system engineering.

Technical level: Advanced. The paper assumes familiarity with quantization formats (NVFP4, INT4, look-up-table encodings), the compute-bound versus memory-bound distinction between prefill and decode, KV caching, GGUF checkpoints, and serving engines such as vLLM and llama.cpp.

Scope: The paper proposes "disaggregated quantization" (DQ), a framework for assigning different quantization formats, weights, and storage locations to the prefill and decode phases of LLM inference, and validates it on instruction-tuned Qwen 3 and Gemma 3 models, released Qwen3.8-27B GGUF decoders, and post-training-quantized models up to 2.8T parameters.

What This Paper Is About

Prefill (processing the prompt) and decode (generating tokens) place opposite demands on hardware: prefill is generally compute-bound at long token counts, while decode at batch one is dominated by weight-loading traffic. Existing quantization treats both phases identically, which forces a single compromise. The paper's goal is to let each phase use the representation that suits it — hardware-native low-precision compute for prefill, compact weight-only encodings for decode — while training the two pathways to work together toward the same generated response.

Key Contributions

  1. Quantization-aware distillation with disaggregation (QADD). A training recipe that uses the SFT assistant-token mask not only for loss masking but also to select which computational pathway each position takes: prompt positions use the prefill pathway, assistant positions use the decode pathway. Both pathways are trained toward a common response objective in one forward-backward pass, and QADD supports either shared or separate master weights, as well as prefill-only adaptation to a frozen decoder.

  2. A three-scheme ladder of disaggregated quantization. (a) Format disaggregation combines quantized-compute prefill with weight-only low-bitwidth decode, using a common set of weights. (b) Full disaggregation trains separate compute-native prefill weights alongside the decode weights. (c) Offloaded disaggregated prefill (ODP) streams the prefill checkpoint from SSD in block-wise buffers that reuse decode weight memory, overlapping loading with compute.

  3. Prefillers for arbitrary frozen weight-only checkpoints. NVFP4 prefill checkpoints trained to augment pre-quantized decoders whose training data, algorithms, or checkpoints are not available, enabling hardware-native prefill without touching the released decode weights.

  4. Large-scale validation. Format disaggregation is also tested through plain post-training quantization, with no retraining, on eight text and multi-modal models up to 2.8T parameters.

Main Findings

  • Phase-specific quantization damage is asymmetric and workload-dependent. On decode-heavy benchmarks, quantizing decode alone to NVFP4 incurs 2–4× the accuracy loss of quantizing prefill alone on most models, and up to 7× on Gemma3-1B. On prefill-heavy tasks, prefill-only quantization incurs 1.1–4.1× the accuracy loss of decode-only quantization on seven of eight models.

  • Format disaggregation (NVFP4 prefill + NVFP4A16 decode) improves decode-heavy accuracy at no added storage or prefill cost. It boosts mean accuracy over uniform NVFP4 by 1.9 and 3.1 points on Qwen 3 and Gemma 3, by 4.5 and 3.7 points for LUT3, and by 2.5 and 1.8 points for LUT2. On prefill-heavy RULER tasks, the effect is under 1.3 points across all model families and formats. Decode is also 2–3% faster than the non-disaggregated scheme by skipping activation quantization.

  • Full disaggregation raises accuracy on both workload types. Decode-heavy gains over non-disaggregated serve are 6.3 and 5.2 points for LUT3 and 10.7 and 7.4 points for LUT2 (Qwen 3 and Gemma 3). Prefill-heavy gains are 4.1 and 4.3 points for LUT3, and 5.3 and 10.5 points for LUT2. For 2-bit models, fully-disaggregated quantization outperforms LUT2A16 weight-only serving by 4.5–12.5 points while also delivering faster prefill.

  • Concrete Qwen 3 / Gemma 3 numbers (Table 1). For Qwen 3, BF16 reaches 66.2 decode-heavy and 84.2 prefill-heavy accuracy at 16.38 GB; 2-bit weight-only drops to 38.4 / 64.1 at 4.66 GB, while 2-bit full disaggregation recovers 45.5 / 76.6 (9.44 GB device), and ODP returns device occupancy to 4.66 GB at 45.5 / 76.6. For Gemma 3, BF16 is 58.6 / 72.0 at 23.53 GB, 2-bit weight-only is 33.7 / 52.4, and 2-bit full disaggregation is 38.2 / 61.3 (11.43 GB), with ODP again restoring 5.38 GB.

  • ODP removes the device-memory cost of the extra prefill checkpoint. On DGX Spark, prefill compute overtakes SSD loading at around 8K context for all Qwen 3 models, and beyond 16K context offloading adds less than 5% latency on Qwen 3 and 8% on Gemma 3. At 16K context the streamed prefill stack is still 1.47× faster than BF16 on Qwen3-8B and 1.58× on Gemma3-12B.

  • Prefillers transform 1-bit GGUF decoders. Using eight freely released Unsloth GGUF checkpoints of Qwen3.8-27B (IQ1_S, IQ1_M, IQ2_XXS, IQ2_S, Q2_K_XL, IQ3_XXS, IQ3_S, Q3_K_XL), training an NVFP4 prefiller improves 1-bit MMLU-Pro accuracy by 32.5 points and MMMU-Pro by 35.3 points, more than doubling weight-only accuracy on both, without modifying the decode checkpoint. Gains of roughly 7.4 and 6.3 points appear at 2-bit decode and shrink at higher bitwidths; at 3-bit decode, the NVFP4 prefiller slightly degrades performance instead. Gains transfer to visual reasoning even though the QADD corpus contains no multi-modal examples.

  • Prefiller speed in llama.cpp. Time to first token drops from 12.27 to 6.90 seconds at 8K context, a 1.78× speedup over the weight-only baseline. Across measured 4K–32K contexts the speedup ranges from 1.38× to 1.78×, with ODP slower at shorter prompts because of SSD loading.

  • Format disaggregation scales to trillions of parameters under PTQ. Applying NVFP4 PTQ without retraining to eight text and multi-modal models up to 2.8T parameters (Qwen 3.8, Gemma 4, Muse Glimmer, Nemotron 3, Kimi-K3), format disaggregation at 4-bit improves point estimates in 11 of 13 model–benchmark combinations, with 6 significant gains at α = 0.05 and no significant degradations.

  • Reported limitation of the accuracy statistics. Error bars describe two standard deviations over temporal averaging within one training run, not uncertainty across independent runs, and Qwen3.8-27B prefiller scores are single complete benchmark evaluations without checkpoint or repeat averaging.

Methodology in Plain English

The authors start from the observation that a transformer linear layer at batch one must load a weight matrix for every single multiply-add, making decode memory-bound, while a long prompt reuses each weight across many tokens, making prefill compute-bound. That asymmetry suggests separate treatments.

They test the sensitivity of each phase by quantizing only prefill or only decode to NVFP4 and watching two families of benchmarks: decode-heavy reasoning tests (GSM8K, MATH-500, MMLU-Pro, with Qwen 3 reasoning both enabled and disabled) and prefill-heavy tests (RULER's 13 long-prompt tasks at 4K, 8K, 16K, and 32K context). Having confirmed that the two phases respond differently, they use these two evaluation families to attribute accuracy changes to each phase.

Training uses QADD, which operates on 100M tokens of the Tülu 3 SFT corpus with a KL-divergence objective against a frozen unquantized BF16 teacher. The key trick is reusing the SFT label mask as a routing signal: tokens with an ignored label (the prompt) take a prefill computation path, assistant tokens take a decode path. Because the loss is computed causally, the final prompt position predicts the first response token, so gradients reach the prefill weights both through that boundary prediction and through the prompt keys and values that decode attends to. Weights and activations are fake-quantized with straight-through estimation while FP32 master weights are updated by AdamW.

For storage, the paper builds a three-rung ladder. Format disaggregation keeps one set of weights and only changes the compute format per phase — for example, NVFP4 prefill with unquantized decode activations. Full disaggregation trains two checkpoints, a native NVFP4 prefill model and a weight-only decode model, each in its own format. ODP then addresses the disk footprint of the second checkpoint: because a prefill block's weights are dead once that block has produced its outputs, ODP loads prefill weights block by block from SSD into two ping-pong buffers, carving that buffer space out of decode weights that are idle during prefill and restoring it before generation.

For the prefiller experiments, the decode side is frozen at its dequantized weights and only the prefill copy of the quantized linear weights, text-stack normalization scales, and a few shared parameters are trained. Latency is measured end-to-end at batch one through vLLM for decode, with prefill transformer-stack latency measured using a custom stack built from vLLM kernels, plus a llama.cpp fork for ODP time-to-first-token.

Why This Matters

Impact on research. The paper reframes quantization as a phase-aware design problem rather than a uniform per-model decision. It shows that format and weight placement can be treated as additional degrees of freedom alongside bitwidth, and it produces a reusable training tool (QADD) that supports jointly optimized or prefill-only adaptation. The finding that the same quantized format damages accuracy differently depending on workload type is a methodological caution for anyone reporting a single aggregate quantization number.

Real-world applications:

  • Local and single-user deployment. ODP lets a second, specialized prefill model live on SSD and be streamed on demand (device memory for 2-bit Qwen 3 stays at 4.66 GB rather than 9.44 GB under full disaggregation), making accurate low-bit models viable on a single device.
  • Accelerating released quantized checkpoints. Prefillers improve already-quantized GGUF decoders — 1-bit Qwen3.8-27B gains 32.5 points on MMLU-Pro — without requiring the decoder to be retrained or its quantization pipeline to be known.
  • Disaggregated datacenter serving. When prefill and decode already run on separate accelerators, each instance can hold a checkpoint in its own format, and the interface between them remains an ordinary KV cache with unchanged layer and head dimensions.
  • Low-bit interactive assistants. The 1.78× time-to-first-token speedup at 8K context in llama.cpp directly targets perceived responsiveness for long-context prompts.

Industry relevance. The results span hardware-native NVFP4 compute, widely used serving stacks (vLLM plus NIXL, llama.cpp with a released ODP fork), GGUF formats released by Unsloth, and production-scale models up to 2.8T parameters from NVIDIA, Google, Meta, Qwen, and Moonshot. The authors note the scheme is mainly a tool for dense LLMs, since mixture-of-experts models have a much worse compute-to-loading ratio.

Future Directions

  • Batched and multi-turn serving. The authors evaluate batch-one decode, prefill-stack timings on DGX Spark, and llama.cpp time-to-first-token, and explicitly do not evaluate highly batched performance, multi-turn, or agentic behavior.
  • Cache-policy robustness. In multi-turn use, cached assistant tokens retain decode-produced KV entries, while rebuilding the same token history through prefill can produce different representations. Whether the model is robust to this discrepancy is untested.
  • Mixture-of-experts and alternative architectures. The paper argues ODP's streaming principle does not transfer cleanly to MoE models because the compute-to-loading ratio grows with the fraction of active parameters; it also does not evaluate combining ODP with separate prefill networks that use learned KV-cache adapters.
  • Benchmark contamination and evaluation scope. The authors report overlap between the 27B training prompts and MMLU-Pro (one exact match, two normalized matches, and three with token-shingle Jaccard similarity of at least 0.3 among 12,032 items) and therefore exclude three benchmarks for that setup — a sign that contamination auditing and broader benchmark coverage remain open work.

Target Audience

Researchers and engineers working on LLM inference efficiency, quantization, and serving systems — particularly those building or operating disaggregated prefill/decode deployments, optimizing low-bit GGUF or NVFP4 checkpoints, or designing quantization-aware training pipelines. It is also relevant to teams running local inference on memory-constrained devices, and less suited to readers without background in quantization formats and inference phase terminology.

Authors’ abstract

Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

Read the original paper