Skip to content
AI.info

Research

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

Overview Research area: Model compression for large language models — specifically weight-only post-training quantization (PTQ) at binary (1-bit) and sub-1-bit precision, plus custom CUDA kernels for

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
arXiv
2602.06694
Published
2026-02-06
Authors
Hyochan Chong, Dongkyu Kim, Changdong Kim, Minseop Choi

AI summary

Overview

Research area: Model compression for large language models — specifically weight-only post-training quantization (PTQ) at binary (1-bit) and sub-1-bit precision, plus custom CUDA kernels for efficient decoding.

Technical level: Advanced. The paper assumes familiarity with quantization (bit-widths, scales, groups), low-rank factorization, ADMM, Hessian/K-FAC preconditioning, and straight-through estimators.

Scope: The paper introduces NanoQuant, a post-training method that factorizes each linear-layer weight into two low-rank binary matrices plus full-precision channel scales, enabling genuine 1-bit and sub-1-bit compression without the metadata overhead that keeps prior binary PTQ methods above 2 bits per weight.

What This Paper Is About

Existing binary PTQ methods binarize weights in place with full-precision scales, which structurally caps them at a minimum of 1 bit per parameter, and they additionally store grouping metadata that pushes their real storage cost above 2 bits — sometimes above 3 bits (for example, STBLLM is reported at 4.13 bits). Binary quantization-aware training (QAT) methods do reach 1-bit and sub-1-bit levels using low-rank binary representations, but they need hundreds of millions or billions of tokens and multiple GPUs over multiple days, which makes compressing 70B-scale models impractical. NanoQuant's goal is to close that gap: a data-efficient, compute-efficient PTQ method that reaches sub-1-bit compression using only a small calibration set on a single GPU.

Key Contributions

  1. A PTQ method reaching both 1-bit and sub-1-bit levels. NanoQuant is described as the only method in the paper's Table 1 that is simultaneously PTQ, scalable to 70B+ models, capable of 1-bit, and capable of sub-1-bit compression — the listed baselines each satisfy only a subset.

  2. A stability analysis of initialization, with empirical evidence that precise low-rank binary initialization is critical. The paper argues that how the binary factors are initialized determines whether sub-1-bit quantization is viable at all, and supports this with an ablation comparing initialization schemes.

  3. Extensive evaluation across model families and tasks. Experiments span Llama-2, Llama-3, Gemma 3, Qwen3, and Rnj-1 at sizes from 0.6B to 70B parameters, measured with WikiText-2 perplexity and zero-shot accuracy on six commonsense reasoning tasks, while using limited calibration data.

  4. Custom binary GEMV and GEMM CUDA kernels. These are implemented to raise inference throughput, reduce memory footprint, and improve energy efficiency across datacenter GPUs, consumer GPUs, and edge devices.

Main Findings

  • Sub-1-bit PTQ is achievable without catastrophic degradation. NanoQuant reports WikiText-2 perplexities at 1.00 bit, 0.80 bits, and 0.55 bits across 17 model configurations. At 1.00 bit, Llama-2-7B reaches 10.34, Llama-2-13B 8.71, Llama-2-70B 6.52, Llama-3-8B 14.97, and Qwen3-8B 12.47, against full-precision values of 5.47, 4.88, 3.32, 6.24, and 7.00 respectively.

  • Prior binary PTQ baselines carry large effective-bit overhead. In Table 4 on Llama-2-7B, BiLLM is listed at 2.88 bits, ARB-LLM_RC at 2.51 bits, HBLLM_R at 3.25 bits, STBLLM at 4.13 bits, and GPTQ (W2g64) at 2.28 bits. NanoQuant is listed at 1.00 bit with a model size of 1.33 GB, versus 13.48 GB for the full-precision model.

  • Competitive zero-shot accuracy at 1 bit. On the six-task average, Llama-3-8B scores 45.95 for NanoQuant at 1.00 bit, versus 39.83 (STBLLM, 4.13 bits), 50.45 (HBLLM_col, 3.25 bits), 38.16 (BiLLM, 2.88 bits), 44.23 (ARB-LLM_RC, 2.51 bits), 36.99 (GPTQ w2g64, 2.28 bits), and 71.44 for BF16. On Qwen3-8B, NanoQuant scores 48.94 at 1.00 bit, against 71.22 for BF16, 46.15 (STBLLM), 54.30 (HBLLM_col), 51.86 (ARB-LLM_RC), and 37.92 (GPTQ).

  • Dramatically lower data and compute than binary QAT. In Table 4, NanoQuant reaches 10.34 perplexity on Llama-2-7B with 0.26M tokens and 1.7 GPU hours, compared with OneBit (155.46M tokens, 700.7 GPU hours, perplexity 9.73), BinaryMoS (196.00M, 92.8, 7.88), DBF (1.38B, 37.6, 9.25), and LittleBit (196.00M, 123.6, 9.08). With 2.10M tokens and 2.5 GPU hours, NanoQuant reports 8.85 perplexity.

  • Initialization choice matters a lot. On Rnj-1 at 0.8 bits, LB-ADMM initialization gives 20.06 perplexity and 39.29 zero-shot accuracy, versus 30.27 / 37.20 for DBF ADMM and 167.73 / 35.11 for Dual-SVID.

  • Every pipeline component contributes. On Qwen3-8B-Base, initialization alone yields 206.03 perplexity; adding error propagation mitigation gives 15.07; adding factorized refinement gives 15.00; combining initialization, error mitigation, and factorized refinement gives 13.58; and adding model reconstruction yields 12.47 perplexity with 48.94 zero-shot accuracy.

  • Perplexity comparable to QAT at far lower cost. In Table 7, NanoQuant on Qwen3-4B-Base uses 1.05M tokens and 2.3 GPU hours for 12.62 perplexity and 50.63 zero-shot, versus LittleBit (169.50M tokens, 92.5 hours, 14.79, 47.32) and DBF (1.19B tokens, 25.3 hours, 14.62, 52.30). On Llama-2-7B, NanoQuant reports 9.01 perplexity / 51.01 zero-shot with 1.05M tokens and 2.1 GPU hours, versus LittleBit at 9.08 / 54.92 and DBF at 9.25 / 54.24.

  • Large compression and consumer-GPU deployment. Llama-2-70B is compressed by 24× — from 137.95 GB to 5.75 GB — in 13 hours on a single H100, and the resulting model runs on a consumer 8 GB GPU at up to 20.11 tokens per second.

  • Inference speed, memory, and energy gains on consumer hardware. On an NVIDIA RTX 3050 (8 GB) with Llama-3.2-1B and 3B, NanoQuant delivers up to 3.6× higher decoding throughput, 5.4× lower peak memory usage, and 3.9× greater energy efficiency compared to BF16 baselines; the speedup figure of 3.6× is stated specifically for Llama-3.2-3B. On an NVIDIA H100 (80 GB) with Llama-2-13B and Qwen3-32B, the paper reports up to 10× lower memory usage during inference versus BF16.

  • A reported Pareto frontier. Figure 6 places NanoQuant on a new efficiency frontier for the Qwen3 family (0.6B, 1.7B, 4B, 8B, 14B) in the low-bit regime.

Methodology in Plain English

NanoQuant treats compressing a weight matrix as a low-rank binary factorization problem. Instead of storing one full-precision weight per parameter, each linear layer's weight is approximated as two binary matrices (entries restricted to −1 or +1) multiplied together, then rescaled by one vector along the output channels and one along the input channels. Because the two binary matrices have an inner rank r that is much smaller than the layer dimensions, the total storage can fall below one bit per original weight. A figure in the paper shows the flow: factorize into continuous latent factors and floating-point scales, binarize the factors, then pack the −1/+1 values into bits (mapping −1 to 0 and +1 to 1) inside 8-bit blocks.

Compression runs block by block through the transformer. For each block, the method does three things:

  1. Mitigate accumulated error. Earlier blocks are already quantized, so their error contaminates the inputs to the current block. NanoQuant fine-tunes the current block's full-precision weights to absorb that error before compressing it. This is applied to all linear layers.

  2. Initialize the binary factors precisely. The paper treats initialization as the make-or-break step for PTQ, where only a small calibration set is available. It first weights the reconstruction objective using second-order information approximated with K-FAC, producing diagonal preconditioners from activation and gradient statistics collected in a global calibration pass. Because those statistics can be skewed by outliers, the diagonal entries are shrunk toward their mean by a coefficient gamma, which the authors find should be around 0.2 for Llama and Qwen models and around 0.6 for Gemma 3 and Rnj-1. It then solves for the factors with an ADMM procedure (LB-ADMM) that alternates between closed-form linear solves for the continuous factors, a sign-preserving rank-1 projection (SVID) for the auxiliary variables, and dual updates. A stabilized Cholesky decomposition is used in the linear solve, which the paper says reduces complexity to O(r³/3) versus O(2r³/3) for general LU factorization — a detail the authors credit for scaling to Llama-2-70B. Finally, because the converged factors have arbitrary relative scale, an equilibrium factor equalizes their Frobenius norms, and the channel scale vectors are extracted as mean absolute values of the balanced rows.

  3. Refine the factorized components locally. The continuous latent factors and the scale vectors are jointly tuned against the full-precision block's outputs using a straight-through estimator, so gradients pass through the sign operation. The sign structure can therefore be adjusted while channel magnitudes are optimized. Once this converges, the signs are taken as the final binary values and packed.

After all blocks are done, a final model-level stage tunes only the floating-point scales, keeping the packed binaries frozen, by minimizing the KL divergence between the quantized model's softened output distribution and the full-precision model's. Freezing the binary weights is what keeps the memory cost low enough to calibrate 70B-scale models on one GPU.

Why This Matters

Impact on research. The paper challenges the assumption that reaching sub-1-bit precision requires expensive end-to-end QAT. It argues that careful initialization plus block-wise reconstruction can substitute for large-scale retraining, and it reframes initialization as a core algorithmic problem rather than preprocessing. It also shows that low-rank binary representations can be a genuine alternative to low-bit integer quantization for memory-critical deployment, and it provides a resource-efficiency comparison table that quantifies how much data and GPU time the QAT route costs.

Real-world applications.

  • Running 70B-class models on consumer 8 GB GPUs or laptops, where the model otherwise would not fit in memory at all.
  • On-device assistants and private, local inference where data cannot leave the device.
  • Edge and embedded deployments, with the paper reporting analysis on an NVIDIA Jetson TX2.
  • Datacenter serving where memory bandwidth, not compute, is the bottleneck, reducing serving cost per token and energy per token.

Industry relevance. The work comes from Samsung Research, targets production inference stacks, and ships code at github.com/SamsungLabs/NanoQuant. It also includes purpose-built binary GEMV and GEMM CUDA kernels, which is what makes the compression translate into measured throughput and energy gains rather than only smaller files.

Future Directions

  • Extending the approach to more architectures and modalities. The current evaluation covers dense decoder-only text models from five families; whether the same low-rank binary factorization holds for mixture-of-experts models (which the paper notes QMoE targets) or non-text models is left open.

  • Reconciling PTQ efficiency with QAT-level accuracy at the lowest bits. At 1.00 bit, NanoQuant is competitive but still trails the strongest binary QAT baselines on zero-shot accuracy in several comparisons; narrowing that gap without adopting QAT's data and compute costs is a natural next step.

  • Pushing below 0.55 bits. The paper reports results at 0.55 bits with clearly degraded perplexity (for example, 16.66 on Llama-2-7B and 32.62 on Rnj-1); how far the low-rank binary factorization can be pushed before it breaks down is not resolved here.

  • Broadening kernel and hardware coverage. The paper references kernel details and Jetson TX2 results in an appendix, and states that optimized inference kernels for the binary PTQ baselines are currently unavailable, which limits head-to-head speed comparisons. Establishing a common kernel benchmark would clarify the real deployment advantage.

Target Audience

This paper is most useful to applied ML engineers and systems researchers who deploy LLMs under tight memory budgets, and to quantization researchers tracking the boundary between post-training and training-based compression. It will also interest hardware and kernel engineers working on binary or low-precision inference, and teams evaluating whether extreme compression can replace expensive retraining pipelines. Readers without background in quantization, ADMM, or Hessian-based reconstruction will find the methodology section demanding; the results tables and the resource-efficiency comparison are accessible to a broader technical audience.

Authors’ abstract

Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binary (1-bit) levels, as they either require large amounts of data and compute or incur additional storage. In this work, we propose NanoQuant, the first post-training quantization (PTQ) method to compress LLMs to both binary and sub-1-bit levels. NanoQuant formulates quantization as a low-rank binary factorization problem, and compresses full-precision weights to low-rank binary matrices and scales. Specifically, it utilizes an efficient alternating direction method of multipliers (ADMM) solver to precisely initialize latent binary matrices and scales, and then tunes the initialized parameters through a block and model reconstruction process. Consequently, NanoQuant establishes a new Pareto frontier in low-memory post-training quantization, and enables sub-1-bit compression. NanoQuant makes large-scale deployment feasible on consumer hardware. For example, it compresses Llama2-70B by 25.8$\times$ in just 13 hours on a single H100, enabling a 70B model to operate on a consumer 8 GB GPU. Code is available at https://github.com/SamsungLabs/NanoQuant.

Read the original paper