Research
CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer for High-Sparsity and Low-Precision Networks
CoDeQ: End-to-End Joint Model Compression with Dead-Zone Quantizer Overview Research area: Neural network model compression — specifically joint pruning and quantization, and quantization-aware traini

- arXiv
- 2512.12981
- Published
- 2025-12-15
- Authors
- Jonathan Wenshøj, Tong Chen, Bob Pepin, Raghavendra Selvan
AI summary
CoDeQ: End-to-End Joint Model Compression with Dead-Zone QuantizerOverview
Research area: Neural network model compression — specifically joint pruning and quantization, and quantization-aware training (QAT). arXiv:2512.12981v1 [cs.LG], published 15 Dec 2025, by Jonathan Wenshøj, Tong Chen, Bob Pepin, and Raghavendra Selvan (Department of Computer Science, University of Copenhagen). Licensed CC BY 4.0.
Technical level: Advanced. The paper is built on formal quantization notation (mid-tread uniform symmetric quantizers, absmax scale factors, straight-through estimators) and derivations of a custom dead-zone quantizer.
Scope: The paper proposes CoDeQ, a fully differentiable method that induces sparsity directly inside a quantizer's dead zone, so that pruning and quantization parameters are learned jointly by backpropagation with a single global sparsity hyperparameter and no auxiliary search procedures.
What This Paper Is About
Pruning and quantization are usually applied one after the other, each with its own fine-tuning stage, which costs training time and can give sub-optimal results because the two decisions are made independently. Existing joint methods try to decide sparsity and bit-widths together, but they typically rely on auxiliary, non-differentiable procedures (discrete searches, ADMM-style projections, evolutionary search, projected optimizers) that sit outside the training loop, add engineering complexity, and give the compression parameters no direct gradient signal. CoDeQ's goal is to make joint pruning–quantization as simple as ordinary quantization-aware training: the pruning decision emerges from the quantizer itself.
Key Contributions
- Unified pruning via quantization. The authors show theoretically that the dead zone of a scalar quantizer is equivalent to magnitude pruning at a threshold of half the dead-zone width, and derive a quantizer with an adjustable dead-zone width that keeps evenly spaced levels in the non-zero region, so no quantization levels are wasted in the pruned region.
- A simple, differentiable joint pruning–quantization method (CoDeQ). The dead-zone width is parameterized and learned by backpropagation during QAT alongside the quantized weights, with sparsity controlled by a single global hyperparameter (plus an optional second one for mixed precision).
- Fixed-bit and mixed-precision support with decoupled control. Bit-width can be fixed by the user while only sparsity is learned, or bit-width itself can be learned; the paper argues this decoupling matters because bit-width is often dictated by hardware and sparsity is the principal remaining degree of freedom.
- Competitive accuracy–compression trade-offs with low engineering overhead. Empirical results across ResNet-20, ResNet-18, ResNet-50, and TinyViT are reported without any auxiliary procedures, and Table 1 reports CoDeQ as needing two hyperparameters, being fully learnable, fixed-bit capable, and architecture agnostic.
Main Findings
- ResNet-20 on CIFAR-10 (Table 2, mean and standard deviation over 3 runs): baseline 32-bit at 91.70% accuracy and 100% relative BOPs. CoDeQ at fixed 4-bit reaches 91.88 ± 0.15% at 2.95 ± 0.12% relative BOPs, and CoDeQ with mixed precision reaches 91.86 ± 0.10% at 2.61 ± 0.08%. The paper states CoDeQ achieves the highest accuracy among compared approaches while also delivering the lowest BOPs, and that the fixed-bit variant outperforms the next-best method, QST, in both accuracy and BOPs.
- ResNet-18 on ImageNet (Table 3, 3 runs): baseline 32-bit at 70.30% and 100% relative BOPs. CoDeQ fixed 4-bit: 69.83 ± 0.14% at 4.75 ± 0.19%; CoDeQ mixed precision: 69.81 ± 0.09% at 4.42 ± 0.05%. For comparison, SQL is reported at 68.60% / 6.20% and QST-B at 69.90% / 5.00%. The abstract describes this as reducing bit operations to roughly 5% while maintaining close-to-full-precision accuracy.
- ResNet-50 on ImageNet (Table 4, one run): baseline 32-bit at 76.30% and 100% relative BOPs. CoDeQ mixed precision reports 75.29% at 2.62% relative BOPs, versus QST-B at 76.10% / 4.50%, GETA at 74.40% / 5.38%, and CLIP-Q at 73.70% / 6.30%. The paper describes this as a small top-1 drop of approximately 0.81% relative to QST while reducing BOPs from 4.5% to 2.62%.
- Vision transformers (Table 5, TinyViT on CIFAR-10, 3 runs): baseline 87.99 ± 0.09% at 100% relative BOPs; CoDeQ fixed 4-bit gives 87.86 ± 0.19% at 2.15 ± 0.04% relative BOPs, and CoDeQ mixed precision gives 87.85 ± 0.21%. The mixed-precision relative BOP entry in the supplied table text is truncated and appears only as "1", so the exact value is not legible in the provided content.
- Fixed-bit and mixed-precision behave nearly identically. The paper repeatedly notes that the fixed-bit models perform close to the mixed-precision models on accuracy and BOPs, which it presents as evidence that CoDeQ is stable in hardware-friendly uniform-precision settings.
- Learned layer-wise patterns match established heuristics. Figure 3 shows that CoDeQ assigns higher precision and lower sparsity to the first and last layers of ResNet-18, and greater sparsity and lower precision to deeper layers.
- Sparsity can arise from two different mechanisms. In the CIFAR-10 ResNet-20 ablation (100 epochs, Figures 4a and 4b), varying the dead-zone regularization λ_dz reduces θ_dz and increases sparsity, while adding weight decay on the weights with λ_dz fixed reaches comparable sparsity by shrinking weights into the zero bin rather than widening it; the paper reports that with weight decay the dead zone remains close even at different sparsity levels.
- Comparison to prior joint methods (Table 1). CoDeQ is listed with 2 hyperparameters and as fully learnable, fixed-bit capable, and architecture agnostic; among the compared methods, GETA and FITC are listed as not fully learnable, and the paper reports that several methods (SQL, QST, CLIPQ, DJPQ, BB) are not architecture agnostic or not fixed-bit capable.
Methodology in Plain English
A standard uniform quantizer rounds weights onto a fixed grid; it already contains a "dead zone" — an interval around zero that maps to the value zero — but in a conventional quantizer that zone is exactly one step wide. The authors' central observation is that this zero bin is functionally identical to magnitude pruning: any weight that lands in it is discarded. So instead of pruning with a mask and then quantizing, they make the dead zone an explicit, tunable quantity.
They define a modified quantizer with two separate knobs: a step size for the non-zero grid, and a dead-zone width. The trick is that these are decoupled — subtracting a correction term inside a ReLU shifts the rounding thresholds around zero — so the zero bin can be widened without changing the spacing of the surviving quantization levels. They also derive a "pruning-aware" scale factor that carves out the dead zone first and then spreads the quantization levels over only the remaining dynamic range, which ensures no levels are wasted in the pruned region and reduces quantization error for the weights that survive. When the dead-zone width equals the step size, this reduces exactly to the ordinary absmax uniform quantizer.
To learn the dead zone, the width is written as a function of a single unbounded trainable parameter passed through a tanh, keeping the width positive and bounded by the weight range. Training then adds an L2 penalty on that parameter to the task loss, which pushes the dead zone wider and therefore prunes more. Gradients flow through the rounding operator via a straight-through estimator; the authors deliberately do not apply straight-through to the sign operators (to avoid distorting gradient magnitudes) but do apply it to the ReLU and the clip, so that already-pruned and saturated weights can still move. A second optional parameter controls a learnable bit-width, mapped through a tanh and rounded into a valid range of bit-widths, with its own L2 penalty.
Experiments fine-tune ResNet-18 and ResNet-50 (TorchVision pretrained) on ImageNet for 120 epochs with cosine annealing, and train ResNet-20 from scratch on CIFAR-10 for 300 epochs; TinyViT is trained from scratch on CIFAR-10 for 200 epochs with AdamW at 3×10⁻⁴ learning rate and batch size 128. Batch sizes are 512 for ResNet-18/20 and 256 for ResNet-50. Quantization is layer-wise (one scale factor per layer), gradients are not tracked through the absmax operation, the 99th quantile of the absmax is used for stability with outliers, and an epsilon of 10⁻⁸ is added to the scale factor to avoid division by zero if an entire layer is pruned. The dead-zone and bit parameters are initialized to 3, giving tanh(3) ≈ 0.995, i.e. a full bit budget and a near-zero dead zone at the start; mixed-precision runs use a minimum of 2 bits and a maximum of 8 bits. Compression is measured in bit operations (BOPs) using an accounting adapted from prior work to reflect unstructured sparsity, with activations left at 32 bits since activations are not quantized.
Why This Matters
Impact on research. The paper reframes pruning as a property of the quantizer rather than a separate operator bolted onto training, and it argues that prior joint methods pay for their compression gains with non-differentiable outer-loop machinery and no direct gradient signal to the compression parameters. If the dead-zone equivalence holds as derived, it gives a clean, fully differentiable route to sparsity and a way to decouple sparsity from bit-width — the latter being a complaint the authors raise about coupling in methods such as QST and DJPQ. The paper also notes that reported layer-wise pruning from at least one prior method concentrates sparsity in early layers, which it suggests may be sub-optimal.
Real-world applications (as motivated by the paper's framing):
- Deploying networks on latency-, memory-, and energy-constrained devices, where the paper notes growth in model scale creates deployment barriers.
- Hardware with native support for uniform 4-bit and 8-bit quantization, since the paper emphasizes that commodity hardware predominantly supports these rather than the non-power-of-two bit-widths common in mixed-precision literature.
- Settings where bit-width is fixed by hardware and sparsity is the remaining knob to turn for efficiency — the paper states this explicitly as motivation for decoupling.
- Transformer-based models as well as convolutional ones, as demonstrated on TinyViT.
Industry relevance. The paper positions CoDeQ as straightforward to implement inside existing QAT and deployment frameworks because it is built on a standard uniform symmetric quantizer, and it reports only two hyperparameters relative to the three-to-five listed for compared methods. However, the paper reports no measured latency, wall-clock speed-up, memory footprint, or energy figures; the efficiency claims are expressed in the paper's BOP metric, which the authors themselves note "reflects arithmetic cost under idealized support for unstructured weight sparsity." All reported pruning results are unstructured.
Future Directions
- Validate the BOP-based efficiency claims on real hardware. The paper acknowledges the metric assumes idealized support for unstructured weight sparsity, and reports no latency or energy measurements; actual speed-ups from unstructured sparsity remain an open question.
- Extend to structured sparsity. All reported results use unstructured pruning, and the authors state they ignore activation sparsity and any structured reduction in input channels in their accounting, so whether the dead-zone formulation can produce channel- or block-structured patterns is untested.
- Resolve the fixed-bit versus mixed-precision gap more systematically. Fixed-bit and mixed-precision CoDeQ perform nearly identically in the reported tables; the paper presents this as a strength, but it invites further study of when the optional bit-width parameter earns its second hyperparameter.
- Understand and control the interaction between dead-zone regularization and weight decay. The ablation shows both mechanisms reach comparable sparsity through different routes, and the authors call this interplay important, but the paper does not provide guidance on choosing between them.
Target Audience
Researchers and practitioners working on model compression, quantization-aware training, and efficient inference — particularly those interested in differentiable, end-to-end formulations of pruning and quantization, and readers evaluating whether joint pruning–quantization methods are practical to adopt. The paper is also relevant to engineers deploying networks under fixed hardware bit-widths, though it assumes fluency with quantization notation and straight-through gradient estimation.
Authors’ abstract
While joint pruning--quantization is theoretically superior to sequential application, current joint methods rely on auxiliary procedures outside the training loop for finding compression parameters. This reliance adds engineering complexity and hyperparameter tuning, while also lacking a direct data-driven gradient signal, which might result in sub-optimal compression. In this paper, we introduce CoDeQ, a simple, fully differentiable method for joint pruning--quantization. Our approach builds on a key observation: the dead-zone of a scalar quantizer is equivalent to magnitude pruning, and can be used to induce sparsity directly within the quantization operator. Concretely, we parameterize the dead-zone width and learn it via backpropagation, alongside the quantization parameters. This design provides explicit control of sparsity, regularized by a single global hyperparameter, while decoupling sparsity selection from bit-width selection. The result is a method for Compression with Dead-zone Quantizer (CoDeQ) that supports both fixed-precision and mixed-precision quantization (controlled by an optional second hyperparameter). It simultaneously determines the sparsity pattern and quantization parameters in a single end-to-end optimization. Consequently, CoDeQ does not require any auxiliary procedures, making the method architecture-agnostic and straightforward to implement. On ImageNet with ResNet-18, CoDeQ reduces bit operations to ~5% while maintaining close to full precision accuracy in both fixed and mixed-precision regimes.