Skip to content
AI.info

Research

Adaptive Fused Prior Transfer for Controllable Generative Image Compression

Adaptive Fused Prior Transfer for Controllable Generative Image Compression Overview Research area: Learned image compression (LIC), specifically controllable generative/perceptual image compression a

Adaptive Fused Prior Transfer for Controllable Generative Image Compression
arXiv
2605.16817
Published
2026-05-16
Authors
Yifei Pei, Ying Liu, Nam Ling

AI summary

Adaptive Fused Prior Transfer for Controllable Generative Image Compression

Overview

Research area: Learned image compression (LIC), specifically controllable generative/perceptual image compression at very low bitrates, combining pretrained codebook-based generative priors with entropy-constrained coding.

Technical level: Advanced. The paper assumes familiarity with nonlinear transform coding, hyperprior entropy modeling, vector-quantized codebooks (VQ-VAE/VQGAN), adversarial training, and rate-distortion-perception tradeoffs.

Scope: This paper introduces AFP-GIC, a controllable generative image codec that transfers an adaptive fused prior from a frozen pretrained AdaCode model into an entropy-constrained compression backbone, using asymmetric encoder-side prior guidance and decoder-side prior prediction.

What This Paper Is About

At very low bitrates — the paper focuses on rates below roughly 0.2 bpp — the compressed representation does not carry enough information to preserve fine textures and local structures, and distortion-oriented reconstruction tends to produce over-smoothed, less natural images. Perceptual and generative codecs address this by synthesizing missing detail with reconstruction priors, and controllable codecs let one model cover many bitrate and fidelity/realism preferences, but controllability alone does not solve the decoder-side prior problem: the decoder still has to infer missing detail from limited transmitted information, and existing controllable designs rely on single-codebook token-based priors. AFP-GIC's goal is to let a controllable generative codec use a richer adaptive fused prior from a frozen pretrained AdaCode model, without transmitting that prior itself and without breaking entropy-constrained decoding.

Key Contributions

  1. Adaptive fused-prior transfer for controllable generative compression. The formulation generalizes prior-guided controllable compression from single-codebook reconstruction priors to a frozen adaptive fused-prior model.

  2. An asymmetric encoder-decoder prior-transfer mechanism. Adaptive fused-prior features guide latent formation at the encoder, while at the decoder a compatible fused prior is predicted from compressed representations and used to guide reconstruction through the frozen pretrained AdaCode decoder, mitigating the prior-availability mismatch in entropy-constrained compression.

  3. An analytical motivation under a Lipschitz decoder assumption, relating improved prior alignment to a tighter reconstruction-error upper bound, and showing that the adaptively fused prior family contains the single-codebook alternative as a special case.

  4. A lightweight fully convolutional decoder paired with the prior transfer, which reduces the parameter footprint and measured decoding latency relative to DC-VIC.

Main Findings

  • Latency reduction: AFP-GIC achieves 18.1% lower decoder latency than DC-VIC.
  • Parameter reduction: AFP-GIC uses 31.10 million (20.5%) fewer inference parameters than DC-VIC.
  • Rate-distortion competitiveness: Experiments on Kodak, CLIC2020, and DIV2K show competitive PSNR and SSIM.
  • Perceptual gains: The clearest perceptual gains are reported in NIQE scores and in very-low-bitrate visual comparisons.
  • Analytical result: Better decoder-side fused-prior alignment tightens a reconstruction-error upper bound under a Lipschitz decoder assumption.
  • Generality of the prior family: The adaptively fused prior family contains single-codebook choices as special cases.
  • Not reported in the provided content: The specific PSNR, SSIM, NIQE, or bpp values on Kodak, CLIC2020, and DIV2K are not given in the supplied text; only the qualitative claim of competitiveness is stated.

Methodology in Plain English

The system starts from an ordinary entropy-constrained learned codec: an encoder turns the image into a latent, a hyperprior and slice-wise context model drive arithmetic coding, and a decoder rebuilds the image. On top of that, AFP-GIC adds a frozen pretrained AdaCode model as a source of generative prior knowledge.

AdaCode's prior is "fused" because, instead of picking one entry from a single codebook at each spatial location, it quantizes the latent with K codebooks and blends the resulting quantized vectors using spatially varying fusion weights α that sum to one across codebooks. That produces a continuous, image-adaptive prior p covering a wider range of structures and textures than a single discrete codebook choice.

The catch is asymmetry. At the encoder, the ground-truth fused prior can be computed directly from the input image, so a lightweight Prior Feature Adapter resizes it to the encoder feature grid and projects it into the encoder channel space, where it is concatenated after the third downsampling layer and added residually. At the decoder, neither the original image nor the ground-truth prior exists. AFP-GIC therefore trains a Prior Estimator — a stack of eight residual blocks with Group Normalization (32 groups) and SiLU activations, an output head, and a skip head — to predict a compatible fused prior from the quantized latent and the control variables.

Reconstruction control comes from two variables, β_rate and β_prior. They are mapped to Fourier-based embeddings, combined by a small MLP, and used to condition both encoder latent formation and, at the decoder, prior prediction plus an SFT (Spatial Feature Transform) extractor that produces scale and shift modulation features. The frozen AdaCode decoder then takes the predicted prior and the SFT features to produce the final image. Because the fused prior is never sent, the bitstream stays limited to the compressed latents plus a compact header index for the selected operating point.

Training proceeds in two phases. First, a non-adversarial dual-conditioned phase optimizes a rate term, an MSE distortion term, an LPIPS perceptual term (AlexNet backbone, also used for the reported LPIPS metric), and a prior-consistency term that pulls the predicted fused prior toward the ground-truth fused prior; the rate and prior-consistency weights are set exponentially via exp(β_rate) and exp(β_prior) so they stay positive and vary smoothly. Second, an adversarial refinement phase freezes the encoder and entropy models, drops the rate term, and adds a conditional PatchGAN discriminator whose real/fake judgment is conditioned on the control pair encoded into eight spatially broadcast channels concatenated with the RGB input. The discriminator is introduced only after the codec has learned stable controllable compression and prior prediction, to avoid destabilizing training.

Why This Matters

Research impact. The paper reframes a practical design question in controllable generative compression: how should reconstruction priors be represented, and how should they be made available to a decoder that has no access to the source image? It argues that moving from a single discrete codebook to a continuous, adaptively fused prior expands the space of usable reconstruction priors, and it offers an analytical argument linking prior alignment to a reconstruction-error bound. The asymmetric encoder/decoder transfer design, plus the observation that single-codebook priors are a special case of the fused family, gives a template that other prior-guided codecs could follow — including swapping in other frozen pretrained generative models.

Real-world applications.

  • Very-low-bitrate image delivery, such as thumbnails, previews, or constrained-bandwidth feeds where bandwidth is the binding constraint.
  • Mobile and edge image pipelines that need one model to serve multiple quality/bitrate operating points instead of several separately trained codecs.
  • Perceptual-quality-sensitive consumer imaging, where visually sharp and natural results matter more than pixel-exact fidelity.
  • Archival or social-media distribution workflows where users can choose a fidelity-versus-realism tradeoff at decode time using a compact operating-point index.

Industry relevance. The measured 18.1% lower decoder latency and 31.10 million (20.5%) fewer inference parameters than DC-VIC speak directly to deployment cost, since decoder compute and memory footprints are often the practical bottleneck for on-device compression. The once-for-all controllability design targets the common industry need to serve many rate and quality targets from a single trained model.

Future Directions

  • Efficiency versus quality tradeoffs in the fused prior. Because the fused-prior mechanism blends K codebooks with predicted weights, an open question is how much of the perceptual gain comes from the fusion itself versus the frozen AdaCode backbone, and where the cost/benefit point lies.
  • Extending the analysis beyond the Lipschitz decoder assumption. The error-bound argument depends on that assumption; whether it holds for the actual learned decoder, and how tight the bound is in practice, is not established.
  • Generalizing the transfer mechanism to other pretrained priors. The paper presents the approach as a way to transfer a frozen AdaCode prior; whether the same asymmetric encoder/decoder design transfers to other frozen generative or codebook-based models is left open.
  • Full reporting and ablation. The provided text reports qualitative dataset claims plus latency and parameter comparisons, but does not include per-dataset bitrate-level details, cross-metric breakdowns by control setting, or ablation results isolating each proposed component; these would clarify exactly where the design pays off.

Target Audience

Researchers and graduate students working on learned image compression, generative and perceptual codecs, and rate-distortion-perception tradeoffs; practitioners building codecs that must serve multiple bitrate or quality targets; and engineers interested in prior-transfer techniques that combine frozen pretrained generative models with entropy-constrained compression pipelines. Readers need a solid background in deep-learning-based compression to follow the entropy-coding, codebook, and adversarial-training details.

Authors’ abstract

At very low bitrates, image compression discards fine textures and local structures, while distortion-oriented reconstruction often produces over-smoothed images. Generative codecs synthesize missing details, but existing codebook-based controllable designs generally rely on single-codebook reconstruction priors. We propose Adaptive Fused Prior Transfer for Controllable Generative Image Compression (AFP-GIC), which transfers an image-adaptive fused prior from a frozen pretrained AdaCode model. Encoder-side prior features guide latent formation, while the decoder predicts a compatible fused prior from the compressed representation and control variables, without transmitting the prior itself. A motivating analysis shows that better decoder-side prior alignment tightens a reconstruction-error upper bound and that the fused-prior family includes single-codebook choices as special cases. A single pretrained model supports five evaluated bitrate operating points. Under the unified benchmark, AFP-GIC achieves 18.1% lower decoder latency and uses 31.10 million (20.5%) fewer inference parameters than DC-VIC. Experiments on Kodak, CLIC2020, and DIV2K show competitive PSNR and SSIM, with the clearest naturalness gains in NIQE scores and very-low-bitrate visual comparisons. Code: https://github.com/yifeipet/AFP_GIC.

Read the original paper