Research
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs Overview Research area: Model compression for large language models — specifically post-training quantization (PTQ) to 1-bit weights,

- arXiv
- 2512.00862
- Published
- 2025-11-30
- Authors
- Ningning Chen, Weicai Ye, Ying Jiang
AI summary
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMsOverview
Research area: Model compression for large language models — specifically post-training quantization (PTQ) to 1-bit weights, combined with signal-processing techniques (wavelet transforms).
Technical level: Advanced. The paper assumes familiarity with quantization theory (salient/non-salient weights, block-wise quantization, Hessian-based importance), transformer linear layers, and frequency-domain transforms.
Scope: A compressed summary of the paper's proposed method (HBLLM) and its reported results on OPT and LLaMA-family models under 1-bit weight quantization.
What This Paper Is About
Running huge language models is expensive because their weights consume enormous memory. "1-bit quantization" tries to shrink every weight down to roughly a single bit, but existing methods either collapse on modern architectures (the paper reports such failures on LLaMA3-8B) or need costly global matrix rotations that slow inference. HBLLM's goal is to push weights to about 1 bit while keeping model quality close to the full-precision (FP16) original, using a cheap, localized Haar wavelet transform plus two structure-aware grouping strategies.
Key Contributions
- A localized orthogonal transformation mechanism: a single Haar wavelet transform decomposes weight matrices into high- and low-frequency components, raising binary expressiveness while cutting transform computation relative to global orthogonal transforms.
- Frequency-aware multi-parameter intra-row grouping: introduces intra-row grouping in the frequency domain to capture structural patterns, splitting each row into frequency bands and further into dense/sparse groups.
- ℓ2-norm-based, saliency-driven column selection: ranks and retains key columns using an ℓ2-norm saliency metric to reduce quantization error.
- Intra-frequency-band mean sharing: for non-salient weights, groups within the same row and wavelet band share a mean, reducing per-parameter storage (the paper states by 0.25 bits) without sacrificing fidelity.
Main Findings
- Perplexity headline result: HBLLM attains a perplexity of 6.71 on LLaMA2-13B with an average weight storage of only 1.08 bits (as stated in the abstract).
- Gap to FP16: Across language modeling tasks (C4, PTB, WikiText2), the perplexity ratio between HBLLM and the original FP16 model stays in the range of 1.2–2.2, outperforming the next-best methods by 33%–66%. Section 4.2 states results of 1.22× to 2.48× the original perplexity.
- Zero-shot accuracy retention: On 9 zero-shot QA benchmarks, HBLLM retains 73.8%–88.8% of the original model's accuracy. Relative QA accuracy gains over prior methods range from −0.73% to +11.3%.
- Stability on modern architectures: On LLaMA3-8B, HBLLM remains stable with no performance collapse, whereas the paper reports other methods degrade sharply (for example, BiLLM at 1.09 W-bits shows perplexity 385.8 on C4 for LLaMA3-70B).
- Expressive capacity metric (CIQ): The paper introduces the cardinality of the Inverse Quantization Set (CIQ). Under 1-bit quantization, CIQ is 8 for BiLLM and 10 for ARB-LLM X; with block size set to 128, ARB-LLM X's CIQ upper bound can reach 128. HBLLM reaches a CIQ of up to 1024 after the Haar transform.
- Compression versus cost: HBLLM-col operates at 1.00 W-bits in the tables, while HBLLM-row ranges around 1.06–1.13 depending on model and size, yet both outperform BiLLM and ARB-LLM RC overall despite lower average bit rates and memory usage in several cases.
- Time overhead: HBLLM increases quantization time by approximately 20%–30% compared to BiLLM across model sizes (44 min vs 36 min at 7B; 98 min vs 71 min at 13B; 173 min vs 142 min at 30B). ARB-LLM X and FrameQuant fail to complete quantization for LLaMA-1-13B and LLaMA-1-30B on a single 24 GB GPU, while HBLLM succeeds.
- Memory: For LLaMA-1, FP16 uses 13.48 GB (7B) and 26.03 GB (13B); HBLLM-col uses 2.67 GB and 5.06 GB, comparable to ARB-LLM RC at 2.83 GB and 5.17 GB.
- Inference latency estimate: Estimated at approximately 31.8% of the FP16 baseline inference time, based on GEMV tests on layers from OPT-175B run on an NVIDIA P100 GPU following the GPTQ benchmark setup (because no existing framework fully supports HBLLM's dequantization).
- Ablations on LLaMA2-7B: ℓ2-norm selection beats ℓ1 (Wiki2 10.52 vs 10.78; PTB 89.23 vs 143.7 for HBLLM-row); row-wise grouping beats global (11.08 vs 16.32; 95.58 vs 1990); shared mean slightly helps HBLLM-row (10.52/89.23 vs 11.08/95.58) and is mixed for HBLLM-col (11.33/150.6 vs 12.02/146.1); and 40 partition candidates was chosen as the default after testing 10, 20, 40, and 80.
Methodology in Plain English
The starting point is a BiLLM-style pipeline that separates "salient" weights (the most important columns) from ordinary ones. HBLLM makes four changes:
- Rank columns by their ℓ2 norm instead of ℓ1 or fixed thresholds. The top-K columns are marked salient via a binary mask and quantized separately.
- Fill the gaps left by removed salient columns using a "FillAvg" step, where each missing column is filled with the average of adjacent non-salient columns, so a row-wise Haar transform can still be applied.
- Apply the Haar wavelet transform to weight rows, producing low-frequency and high-frequency coefficients. The Haar transform is implemented with two fixed 1D convolution kernels of size 2 ([1/2, 1/2] and [1/2, −1/2]), giving O(d) cost versus O(d²) for global transforms such as FrameQuant, and it requires no training or storage.
- Quantize group-wise with sign-based binarization centered on each group's mean, using frequency-aware intra-row grouping (dense and sparse groups per frequency band) and sharing a single mean across the two groups in a band within each row.
The paper presents a Frobenius-norm reconstruction objective but explicitly states it does not solve this objective by explicit optimization; the formulation is a conceptual guide, and the actual method is a set of heuristics and structure-aware strategies approximating it efficiently. GPTQ-style block-wise error correction is used at the layer level.
Experimental setup: Evaluations use PyTorch on NVIDIA GeForce RTX 3090 GPUs with 24 GB of memory. Calibration uses 128 samples from C4 with a sequence length of 2048, and block size 128 for BiLLM, PB-LLM, ARB-LLM, and HBLLM. Models include OPT 1.3B/2.7B/6.7B/13B/30B, LLaMA-1 and LLaMA-2 at 7B/13B/30B/65B/70B, LLaMA-3 8B/70B, and f-R1-Distill-Llama-8B. Perplexity is measured on C4, WikiText2, and PTB; zero-shot accuracy uses PIQA, BoolQ, OpenBookQA, WinoGrande, ARC-e, ARC-c, HellaSwag, COPA, and LAMBADA via LM-Evaluation-Harness. Baselines are BiLLM, ARB-LLM (X and RC variants), PB-LLM (salient weight ratio set to 10% to keep average bit width below 2 bits), and FrameQuant with redundancy factors r=1.0 and r=1.1.
Why This Matters
Impact on research: HBLLM challenges the assumption that global orthogonal rotations are necessary for high-fidelity low-bit quantization, showing a localized wavelet transform can deliver stronger expressiveness (measured by CIQ) at far lower computational complexity. It also introduces CIQ as a theoretical lens for comparing quantization methods, and the paper reports competitive results against QAT-style approaches.
Real-world applications:
- On-device LLM inference where memory and power budgets are tight (the paper describes edge deployment as the target).
- Serving large models in low-resource environments where full-precision weights cannot fit.
- Reducing memory footprint of inference servers to increase model density per accelerator.
- Deployment scenarios such as the NVIDIA P100 setting used for its latency estimation.
Industry relevance: The paper reports that HBLLM quantizes LLaMA-1 up to 30B on a single 24 GB GPU where ARB-LLM X and FrameQuant cannot complete the job, and that it estimates inference at roughly 31.8% of FP16 latency. Both attributes relate directly to practical cost of ownership of LLM serving.
Future Directions
- Extension to Mixture-of-Experts models: The conclusion states that current HBLLM supports only quantized dense models and that the authors will next focus on a MoE PTQ algorithm.
- Direct support in inference frameworks: The latency number is an estimate because no existing inference framework fully supports HBLLM's dequantization; building that support would let measured end-to-end latency be reported instead of estimated.
- Improving HBLLM-col fidelity: The paper notes HBLLM-col trades data fidelity for storage advantages relative to HBLLM-row, pointing to a possible accuracy-versus-storage optimization target.
- Scaling and generalization testing: The results cover OPT, LLaMA-1/2/3, and f-R1-Distill-Llama-8B; whether the gains hold on other architecture families is not reported.
Target Audience
Researchers and engineers working on LLM compression, low-bit quantization, and efficient inference, particularly those already familiar with BiLLM, GPTQ, or ARB-LLM. It also suits practitioners who need to deploy large models under strict memory limits, and readers interested in applying signal-processing tools such as wavelet transforms to neural network compression. Beginners would find the notation-heavy method section demanding but could follow the motivation and results sections.
Authors’ abstract
We introduce HBLLM, a wavelet-enhanced high-fidelity $1$-bit post-training quantization method for Large Language Models (LLMs). By leveraging Haar wavelet transforms to enhance expressive capacity through frequency decomposition, HBLLM significantly improves quantization fidelity while maintaining minimal overhead. This approach features two innovative structure-aware grouping strategies: (1) frequency-aware multi-parameter intra-row grouping and (2) $\ell_2$-norm-based saliency-driven column selection. For non-salient weights, a shared mean is employed across quantization groups within each frequency band to optimize storage efficiency. Experiments conducted on the OPT and LLaMA models demonstrate that HBLLM achieves state-of-the-art performance in $1$-bit quantization, attaining a perplexity of $6.71$ on LLaMA$2$-$13$B with an average weight storage of only $1.08$ bits. Code available at: https://github.com/Yeyke/HBLLM.