Research
Bridging Degradation Discrimination and Generation for Universal Image Restoration
Overview Research area: Low-level computer vision, specifically universal (all-in-one) image restoration and real-world super-resolution using diffusion models. Technical level: Advanced. The paper as

- arXiv
- 2602.00579
- Published
- 2026-01-31
- Authors
- JiaKui Hu, Zhengjian Yao, Lujia Jin, Yanye Lu
AI summary
Overview
Research area: Low-level computer vision, specifically universal (all-in-one) image restoration and real-world super-resolution using diffusion models.
Technical level: Advanced. The paper assumes familiarity with denoising diffusion probabilistic models, latent diffusion, gray-level co-occurrence matrices, and standard restoration benchmarks.
Scope: The paper proposes a single framework, BDG (Bridging Degradation discrimination and Generation), that couples a fine-grained degradation descriptor (MAS-GLCM) with a pre-trained diffusion generator through a three-stage training schedule, and evaluates it on 5D all-in-one restoration, mixed/composited degradation, and real-world super-resolution.
What This Paper Is About
Universal image restoration requires one model to remove many kinds of degradation while still producing detailed, faithful images. The authors argue that existing methods fall into two camps that each solve only half the problem: degradation-discrimination methods (which identify what is wrong with the input) tend to produce over-smoothed outputs, while diffusion/generative-prior methods (which excel at texture) can misinterpret mildly degraded inputs as heavily degraded and hallucinate detail inconsistent with the low-quality image. BDG's goal is to give one model both capabilities: accurate discrimination of degradation type and severity, and the generative prior needed for rich texture, without changing the diffusion architecture.
Key Contributions
-
MAS-GLCM (Multi-Angle and multi-Scale Gray Level Co-occurrence Matrix): a degradation characterization built by averaging standard GLCM matrices computed over several distances (scales) and orientations (angles), designed to capture degradation-specific information while discarding image-content information.
-
Evidence that MAS-GLCM discriminates degradation finely: through visualization across five degradation levels for downsampling, Gaussian noise, and low-light; T-SNE clustering with lower KL divergence than low-quality images; and KNN classification of degradation type and level.
-
A three-stage diffusion training scheme (generation, bridging, restoration) derived by modifying the coefficients in the diffusion backward equation, so the model retains its generative prior while acquiring degradation-discrimination ability.
-
Alignment of MAS-GLCM features with diffusion features in the bridging and restoration stages, plus a "full negative contrastive learning" loss used only in the restoration fine-tuning stage for the real-world super-resolution setting.
Main Findings
-
Degradation classification: Using a KNN classifier on haze, low light, snow, Gaussian noise, and Pepper noise (with Gaussian noise variances 15, 25, 50, 75, 100), MAS-GLCM reaches 97.13% accuracy on degradation type and 74.17% on degradation level. LQ images score 51.44% / 20.00%, Sobel 40.80% / 23.33%, Laplace 83.05% / 20.83%, and Fourier 65.80% / 30.83%.
-
5D all-in-one restoration (Table 2): BDG attains the best score on every reported task — deraining 34.75 PSNR / 0.974 SSIM, enhancement 27.42 / 0.930, desnowing 32.86 / 0.950, dehazing 34.33 / 0.993, deblurring 31.11 / 0.904. Against DiffUIR-L, the gains are 3.72 dB (deraining), 2.30 dB (low light enhancement), 1.39 dB (dehazing), and 1.94 dB (deblurring). Against DCPT-PromptIR, deraining improves by 2.46 dB.
-
Real-world all-in-one, zero-shot (Table 3): BDG records the lowest PIQE in low-light enhancement (34.44) and PIQE/BRISQUE of 31.45 / 24.00 for snow and 47.59 / 34.75 for haze, achieving the majority of best and second-best results among DA-CLIP, InstructIR, DCPT-NAFNet, UniRestore, and FoundIR.
-
Mixed degradation (Table 4): BDG improves over the previous state of the art across all composited scenarios, with the largest gain of 4.28 dB in the haze-plus-rain (H+R) setting.
-
Real-world super-resolution (Table 5): BDG scores the highest or second highest PSNR, SSIM, and LPIPS on all three benchmarks. On DIV2K-Val it reaches 24.1977 PSNR and 0.6241 SSIM, outperforming the second-best diffusion method by 2.45 dB in PSNR; it also leads on DrealSR (28.7961 PSNR) and achieves 0.8039 SSIM there. Non-reference metrics (MANIQA, MUSIQ, CLIPIQA) are competitive rather than dominant — the authors state BDG outperforms its Stable Diffusion 2 baseline and ResShift on those metrics, but it trails SeeSR and DiffBIR on MANIQA, MUSIQ, and CLIPIQA for DIV2K-Val.
-
Ablation, training stages (Table 6): For all-in-one restoration, 300k bridging iterations and 0 RFT gives 30.25 PSNR / 0.871 SSIM; 0 bridging and 300k RFT gives 31.03 / 0.908; 150k of each gives 32.09 / 0.950. The same ordering holds for super-resolution (28.35 / 0.7988 / 0.5839; 28.73 / 0.8093 / 0.4787; 28.80 / 0.8039 / 0.6053).
-
Ablation, losses (Table 7): Removing the degradation-classification loss causes a large drop — 20.88 PSNR / 0.847 SSIM versus 32.09 / 0.950 with all losses — which the authors attribute to collapse of the MAS-GLCM encoder and consequent misalignment of diffusion features. In the restoration fine-tuning stage, using both the bridge loss and the full-negative contrastive loss gives 28.80 / 0.8039 / 0.6053, the best of the four tested combinations.
-
Real-world framing of degradation: Because real degradation cannot easily be assigned discrete class labels, the authors replace type classification with "order classification" over eight intermediate degradation states produced by the Real-ESRGAN-style chain of blur, downsampling, JPEG compression, sinc artifacts, and noise.
Methodology in Plain English
The authors start from the observation that a standard GLCM — a matrix counting how often pairs of gray values appear at a fixed offset — throws away image content and keeps texture/statistics information. They average GLCMs computed over several offsets (different lengths and angles) to get MAS-GLCM, which they show varies strongly with degradation level and clusters degradation types well.
On the generative side, they write the diffusion forward process as an update involving three coefficients: one multiplying the residual (low-quality minus high-quality image), one multiplying Gaussian noise, and one multiplying the low-quality image itself. By choosing which coefficients are zero, they get three regimes. In the generation stage both the residual and low-quality coefficients are zero, leaving a formula formally equivalent to a variance-exploding SDE denoising step, so the model learns only the image distribution. In the bridging stage the residual coefficient is turned on and the low-quality coefficient stays off, so the model sees degradation-relevant residual information while keeping its generative prior; here the features of a MAS-GLCM encoder are aligned with the first-layer decoder features of the UNet via a cross-entropy similarity loss, and an MLP head applies a degradation-classification loss to keep the MAS-GLCM features discriminative. In the restoration stage all coefficients are active, and the low-quality image is injected directly, which the authors say gives stronger fidelity; sampling predicts both noise and residual, unlike DiffUIR, which they say predicts only residual and therefore forgoes the generative prior.
Training proceeds through these three stages in sequence (generation pre-training, bridging, restoration fine-tuning). The bridging and restoration stages each last 150k iterations of a 300k-iteration total. Defaults are batch size 256, learning rate 3×10⁻⁴, AdamW with (β₁, β₂) = (0.9, 0.95), and a loss-balancing weight λ of 0.1. For super-resolution, an extra contrastive loss pushes apart MAS-GLCM features of different degraded samples; the authors deliberately do not apply it during the bridging stage because the feature extractor is still training and they say it would cause representation collapse. The diffusion architecture is not modified — experiments use a 36M UNet pre-trained on ImageNet for all-in-one and mixed degradation, and Stable Diffusion 2 as the super-resolution baseline.
Why This Matters
Research impact: The paper argues that fidelity and perceptual quality need not be traded off in diffusion-based restoration. The claimed mechanism — aligning an interpretable, content-independent degradation descriptor with intermediate diffusion features — is a general recipe that could be applied to other conditional generation problems where the condition carries degradation or quality information. The reported PSNR advantage on full-reference metrics, where diffusion methods often underperform, is the paper's central empirical claim.
Real-world applications (from the paper's task list):
- Removing rain, snow, haze, blur, and low-light degradation from photographs in a single deployed model.
- Restoring images degraded by real, unknown, and composited pipelines (blur + downsampling + JPEG compression + sinc artifacts + noise), i.e., images that have passed through multiple real processing steps.
- Real-world super-resolution of images from synthetic and real benchmarks such as DIV2K-Val, DrealSR, and RealSR.
- Zero-shot handling of degradation types and levels not seen during training, which the real-world experiments test.
Industry relevance: A single model that covers many degradations reduces the need to train and ship separate task-specific models, and because BDG leaves the underlying diffusion architecture unchanged, it can be dropped onto an existing pre-trained generator rather than requiring a new one. The authors explicitly note in their ethics statement that the technology should not be misused for forging misleading images or for malicious restoration/enhancement, which is a direct industry-deployment concern.
Future Directions
- Closing the gap on non-reference perceptual metrics. BDG leads on full-reference metrics but trails methods such as SeeSR and DiffBIR on MANIQA, MUSIQ, and CLIPIQA, so further work could reconcile the two metric families.
- Extending beyond the three evaluated tasks. The paper reports evaluation only for all-in-one restoration, mixed degradation, and real-world super-resolution; other restoration tasks are not covered in the provided content.
- Handling degradation class labels in the wild. The authors replace type classification with an eight-level "order" surrogate for super-resolution because real degradation is hard to label; whether finer or learned pseudo-labelings improve results is left open.
- Understanding and preventing the reported collapse. The ablation shows removing the degradation-classification loss drops all-in-one PSNR from 32.09 to 20.88, and the authors attribute this to MAS-GLCM encoder collapse; the conditions under which the alignment loss may still hurt are not fully characterized here.
Target Audience
Researchers and graduate students working on image restoration, diffusion-based image synthesis, and low-level vision benchmarks; engineers building production image enhancement or upscaling pipelines who need one model to cover many degradation types; and readers interested in combining interpretable handcrafted descriptors (GLCM statistics) with large pre-trained generative models. Readers will need working knowledge of diffusion sampling equations and standard restoration metrics (PSNR, SSIM, LPIPS, FID, PIQE, BRISQUE, MANIQA, MUSIQ, CLIPIQA) to follow the experimental sections closely.
Authors’ abstract
Universal image restoration is a critical task in low-level vision, requiring the model to remove various degradations from low-quality images to produce clean images with rich detail. The challenges lie in sampling the distribution of high-quality images and adjusting the outputs on the basis of the degradation. This paper presents a novel approach, Bridging Degradation discrimination and Generation (BDG), which aims to address these challenges concurrently. First, we propose the Multi-Angle and multi-Scale Gray Level Co-occurrence Matrix (MAS-GLCM) and demonstrate its effectiveness in performing fine-grained discrimination of degradation types and levels. Subsequently, we divide the diffusion training process into three distinct stages: generation, bridging, and restoration. The objective is to preserve the diffusion model's capability of restoring rich textures while simultaneously integrating the discriminative information from the MAS-GLCM into the restoration process. This enhances its proficiency in addressing multi-task and multi-degraded scenarios. Without changing the architecture, BDG achieves significant performance gains in all-in-one restoration and real-world super-resolution tasks, primarily evidenced by substantial improvements in fidelity without compromising perceptual quality. The code and pretrained models are provided in https://github.com/MILab-PKU/BDG.