Skip to content
AI.info

Research

PreResQ-R1: Response-Preference Disentangled Ranking-and-Scoring Reinforcement Optimization for Robust Visual Quality Assessment

Overview Research area: No-reference visual quality assessment (QA) using multimodal large language models (MLLMs), covering both image quality assessment (IQA) and video quality assessment (VQA). Tec

PreResQ-R1: Response-Preference Disentangled Ranking-and-Scoring Reinforcement Optimization for Robust Visual Quality Assessment
arXiv
2511.05393
Published
2025-11-07
Authors
Zehui Feng, Weichuan Wang, Xiaohan Chen, Ting Han

AI summary

Overview

  • Research area: No-reference visual quality assessment (QA) using multimodal large language models (MLLMs), covering both image quality assessment (IQA) and video quality assessment (VQA).
  • Technical level: Advanced. The paper assumes familiarity with reinforcement learning for language models (GRPO, policy gradients), chain-of-thought reasoning, reward modeling, and standard IQA/VQA evaluation metrics (SRCC, PLCC).
  • Scope in one sentence: The paper proposes PreResQ-R1, a reinforcement-optimization framework that separates response-level stability from preference-level ranking/scoring alignment, and evaluates it zero-shot across ten IQA datasets and five VQA datasets.

What This Paper Is About

No-reference quality assessment must predict how humans judge visual fidelity without access to a pristine reference image, and modern MLLM-based systems try to do this by predicting both an absolute score and a relative ranking of quality. When score regression and ranking are optimized inside a single unified objective, and when the model decodes stochastically, the predictions become unstable and generalize poorly across datasets. The paper argues this instability is amplified by response variability in MLLM reasoning, and its goal is to improve robustness by explicitly disentangling response stabilization from preference learning.

Key Contributions

  1. A disentangled reinforcement learning formulation that models response-level stability and preference-level alignment separately, using complementary ranking-aware rewards.
  2. A fine-grained reward scheme that jointly optimizes ranking consistency and scoring alignment through pairwise and triplet ranking signals, including a weak intra-sample response-ranking reward and a magnitude-aware alignment term.
  3. A temporal-spatial evidence aggregation strategy for video QA that avoids dense motion modeling or architectural changes, sampling M frames to form a temporal-flow image and N frames to form local spatial-flow images.
  4. Extensive zero-shot experiments across diverse IQA and VQA benchmarks, reported alongside qualitative reasoning traces, ablation studies on reward components, computational cost analysis, and training-dynamics analysis.

Main Findings

  • IQA benchmark performance: Training only on KADID-10K, PreResQ-R1 reaches an average SRCC of 0.811 and PLCC of 0.795 across 10 IQA benchmarks, described as surpassing prior approaches by 5.6% and 2.2% respectively. The abstract states an improvement over comparable methods of 5.60% on IQA benchmarks.
  • Highest SRCC on every IQA dataset: The paper reports that PreResQ-R1 attains the best SRCC on all evaluated IQA datasets.
  • VQA benchmark performance: PreResQ-R1 achieves an average SRCC of 0.872 and PLCC of 0.891 across 4 datasets. The abstract reports a 2.53% improvement on VQA benchmarks.
  • Transfer from IQA to VQA: The VQA model (PreResVQA-R1) is fine-tuned from the IQA model in a single-prompt stage, showing that IQA perceptual priors transfer to video without extra video-specific encoding parameters.
  • Reward ablation gains: Pairwise preference learning (PSR) alone gives a stable baseline; adding the response triplet ranking reward (RTR) yields roughly a 3% gain; adding the preference triplet ranking reward (PTR) yields a further 3% improvement. The full combination of RTR, PSR and PTR produces the best results in the ablation table.
  • Response/preference dimension ablation: Compared with a baseline of 0.708 SRCC and 0.734 PLCC, using both response and preference dimensions in both training stages reaches 0.811 SRCC (+0.103) and 0.790 PLCC (+0.056).
  • Exploration-to-stability fine-tuning: ESS is reported to accelerate convergence and improve multi-granularity prediction accuracy by transitioning smoothly from exploration to stable optimization; Stage 1 benefits from high response diversity, while a balanced generation-comparison scheme in Stage 2 stabilizes preference learning.
  • Training dynamics: Reward standard deviation moves away from a single static target, CoT score standard deviation evolves low-to-high then reduces slightly in Stage 2, and the RTR reward further stabilizes response generation for faster convergence.
  • Computational cost: The full IQA configuration (RTR+PSR+PTR) uses 71.28 ± 2.03 GB training memory and 2.38 ± 0.23 h/10K samples, with test memory 33.61 ± 0.83 GB and test time 0.72 s/sample. Adding the temporal-spatial flow (TSF) strategy for VQA raises training memory to 73.37 ± 2.91 GB and training time to 10.21 ± 0.38 h/10K samples, with test time 2.95 s/sample. VQA-Thinker is listed at 42.62 ± 0.61 GB and 3.83 s/sample. The text states IQA inference time is 0.6 s/sample and that TSF remains substantially lighter than full video encoding.
  • FineVD distortion categories: PreResQ outperforms Q-Align, VisualQuality and VQAThinker on all five distortion categories (Color, Noise, Artifact, Blur, Temporal); for example, its Temporal SRCC/PLCC is 0.595/0.573 versus 0.539/0.561 for VQAThinker.
  • KonIQ-trained results and hyperparameter sweep: Results for training on KonIQ are reported in an appendix table, where PreResQ-R1 reaches an average of 0.767 SRCC and 0.762 PLCC. A VQA hyperparameter sweep peaks at M = 12 temporal-flow frames and N = 3 spatial-flow frames.

Methodology in Plain English

The researchers start from a pretrained Qwen2.5-VL-7B model and fine-tune it with GRPO, a reinforcement learning method that compares multiple sampled answers per input. Each sample is run through the model K times to produce a chain-of-thought plus a score between 1 and 5. For images, the chain-of-thought explicitly reasons about five quality aspects: saturation, granularity and sharpness (pixel-level), and foreground and background (semantic-level).

Rather than using one scalar reward, the reward signal is split into a response dimension and a preference dimension. The response dimension compares a generation's scores to the median of a triplet of stochastic generations for the same input, which suppresses wild outliers while keeping meaningful variation. The preference dimension compares generations across different samples: a pairwise term checks whether the predicted ordering matches the ground-truth MOS ordering at the same rank index and adds a magnitude-aware term that rewards closeness between predicted and ground-truth score gaps, while a triplet term enforces transitivity so the global ordering stays consistent. A format reward keeps the answer structure valid, and the components are combined with weights alpha and beta.

Training proceeds in two stages designed to move from exploration to stability. In the first stage the model is given X = 5 randomly sampled prompts per sample and K = 12 generations, plus a standard-deviation penalty on the five-dimensional score vector when the spread falls below a threshold, which encourages diverse reasoning early on. The second stage reduces to K = 6 generations and relies on the preference rewards. Settings are alpha = 0.5 and beta = 0.75. Optimization uses AdamW with an initial learning rate of 3 x 10^-7 and linear decay, on 8 NVIDIA A800 GPUs with per-GPU batch size 48 for IQA and 25 for VQA, for 8 epochs on IQA and 1 epoch on VQA.

For video, instead of modeling motion densely, the model uniformly samples M frames into a temporal-flow image that summarizes long-range evolution and N frames into local spatial-flow images for frame-level detail. The video model is initialized from the image model and trained with the same reward structure.

Why This Matters

  • Research impact: The paper reframes QA stability as a property of response-level reasoning consistency rather than purely an optimization artifact, and offers a concrete disentangled reward design that other reasoning-based assessment tasks could adopt.
  • Real-world applications:
    • Automatic quality filtering of user-uploaded images and videos on sharing platforms.
    • Monitoring and tuning image enhancement, restoration or generation pipelines, where the paper frames QA as fundamental to image enhancement, image generation and computational photography.
    • Video streaming and UGC platforms that need no-reference quality scoring of in-the-wild distorted content.
    • Evaluating AI-generated imagery, since AGIQA-3K is among the evaluated benchmarks.
  • Industry relevance: The framework is built on an open 7B multimodal model, requires only synthetic KADID-10K and KonIQ for IQA and LSVQ-28K for VQA, and the paper reports parameter-efficient transfer from IQA to VQA. The conclusion notes it achieves comparable performance with only 6K image samples and 28K video samples, which matters for teams with limited labeled quality data.

Future Directions

  • Develop more principled formulations of perceptual uncertainty, including unified models that jointly capture stability and preference under stochastic generation, since the paper concedes the clean separation of the two dimensions may not hold in human perceptual judgment.
  • Extend the framework to long-form and dynamic visual content with richer temporal modeling, because the current video extension relies on sparse temporal evidence and may be insufficient for fine-grained motion reasoning.
  • Improve robustness under distribution shifts and highly subjective or out-of-distribution conditions where quality perception has no well-defined consensus.
  • Incorporate adaptive or human-in-the-loop feedback, and integrate image quality assessment with text-to-image generation in a unified framework for complementary optimization.

Target Audience

Researchers and graduate students working on no-reference image and video quality assessment, multimodal LLM reasoning, and reinforcement learning from preference or reward signals. It is also relevant to practitioners who need robust, interpretable quality scoring models built on open-weight multimodal backbones, and to engineers evaluating the compute trade-offs of reward design and video evidence aggregation.

Authors’ abstract

Visual Quality Assessment (QA) seeks to predict human perceptual judgments of visual fidelity. While recent multimodal large language models (MLLMs) show promise in reasoning about image and video quality, existing approaches mainly rely on supervised fine-tuning or rank-only objectives, resulting in shallow reasoning, poor score calibration, and limited cross-domain generalization. We propose PreResQ-R1, a Preference-Response Disentangled Reinforcement Learning framework that unifies absolute score regression and relative ranking consistency within a single reasoning-driven optimization scheme. Unlike prior QA methods, PreResQ-R1 introduces a dual-branch reward formulation that separately models intra-sample response coherence and inter-sample preference alignment, optimized via Group Relative Policy Optimization (GRPO). This design encourages fine-grained, stable, and interpretable chain-of-thought reasoning about perceptual quality. To extend beyond static imagery, we further design a global-temporal and local-spatial data flow strategy for Video Quality Assessment. Remarkably, with reinforcement fine-tuning on only 6K images and 28K videos, PreResQ-R1 achieves state-of-the-art results across 10 IQA and 5 VQA benchmarks under both SRCC and PLCC metrics, surpassing by margins of 5.30% and textbf2.15% in IQA task, respectively. Beyond quantitative gains, it produces human-aligned reasoning traces that reveal the perceptual cues underlying quality judgments.

Read the original paper