Skip to content
AI.info

Research

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

Overview Research area: Computer vision / multimodal evaluation — automatic quality assessment of AI-generated and AI-edited images, combining scalar scoring with spatial defect localization. Technica

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations
arXiv
2610.00994
Published
2026-10-01
Authors
Xianda Du, Max Ku, Weiming Ren, Zhi Rui Tam, Chunlin Ren, Ping Nie, Min-Hung Chen, Wenhu Chen

AI summary

Overview

Research area: Computer vision / multimodal evaluation — automatic quality assessment of AI-generated and AI-edited images, combining scalar scoring with spatial defect localization.

Technical level: Advanced. The paper assumes familiarity with vision-language models (VLMs), supervised fine-tuning, GRPO-style reinforcement learning, and ranking-correlation metrics such as SRCC.

Scope: The paper introduces VIEScore2, a single VLM-based evaluator that jointly predicts quality scores and spatially grounded defect locations for image generation and editing tasks, trained on 38K examples from five data sources.

What This Paper Is About

Existing automated image evaluators usually return only a scalar quality score and give no indication of which parts of an image drove that score. The authors build VIEScore2, a unified evaluator that takes a generated image, its prompt, optional conditioning images, and an evaluation instruction, then predicts both quality scores and defect locations in one model pass. The defects are written out as text-native coordinates on a 16 × 16 grid, which a parameter-free parser converts into readable, spatially grounded explanations.

Key Contributions

  1. Unified scoring, localization, and explanation. A single evaluator handles image generation and editing with zero, one, or multiple conditioning images, producing quality scores and defect locations without task-specific prediction heads. A parameter-free parser turns the structured predictions into readable feedback.
  2. Verifiable spatial post-training. Defects are represented as sparse cells on an image grid, a format that is directly parseable and supports a coverage-weighted cell-level Dice reward that explicitly rewards localization overlap.
  3. Learning across heterogeneous sources. Score and spatial annotations from five data sources (ImagenWorld, RichHF-18K, EvalMuse-40K, PAL4VST, COCO) are normalized into a common training format, keeping only the supervision each example actually provides, spanning score-only, localization-only, and joint supervision.
  4. GRPO post-training with combined rewards. Starting from an SFT checkpoint of Qwen3-VL-8B, GRPO with rewards combining cell-level Dice, score accuracy, and output-format validity is shown to improve localization beyond additional supervised fine-tuning alone.

Main Findings

  • Score correlation lead on the primary suite: VIEScore2 reaches an overall-score SRCC of 0.601 across 900 examples with matched conditioning-image inputs, versus 0.491 for Gemini-3-Flash, 0.437 for GPT-5.6-sol, 0.402 for GPT-5.6-terra, and 0.373 for an unfine-tuned Qwen3-VL-8B judge.
  • Joint evaluation advantage: On the joint scoring-and-localization comparison, VIEScore2 records grid IoU 0.324, F1 0.506, PQ SRCC 0.558, and SC SRCC 0.564, compared with Qwen3-VL-8B (0.070 / 0.273 / 0.181 / 0.352) and GPT-5.6-terra (0.112 / 0.247 / 0.357 / 0.508, with valid scores for 554/700 examples).
  • Localization ranking: VIEScore2 ranks first in per-image grid IoU on RichHF (0.299), PAL4VST (0.335), and SynthScars (0.234), second on HAD, and third on AbHuman — top-three on five of six benchmarks, including datasets beyond its training sources. It also leads SynthScars F1 (0.368).
  • Other methods lead elsewhere: GPT-5.6-sol leads HAD IoU, Claude Opus 5.5 leads AbHuman F1, and SDG leads on SDG-30K, indicating no single method dominates all six benchmarks.
  • Task-level variation in scoring: VIEScore2 leads on RichHF, EvalMuse, and SRIG, while Gemini-3-Flash leads on TIE and SRIE, GPT-5.6-sol on MRIG and MRIE, and GPT-5.6-terra on TIG. The paper also notes VIEScore2 shows competitive text-to-image performance on the MMRB2 preference evaluation but weaker editing accuracy than the general-purpose VLM baselines.
  • Conditioning images help: On 250 conditional examples, providing the available conditioning images improves overall-score SRCC from 0.420 to 0.451 and localization F1 from 0.514 to 0.536.
  • GRPO beats extra SFT: Across three independent runs, GRPO improves grid IoU by 0.024 to 0.029 over the shared SFT checkpoint, whereas an additional SFT epoch yields only 0.001.
  • The Dice reward carries the localization gain: A Dice-only variant raises grid IoU from 0.296 to 0.320, while the full reward reaches 0.324 IoU and 0.506 F1 at overall-score SRCC 0.601. The score reward does not measurably change scoring on the primary suite; its main effect is an F1 gain from 0.481 to 0.506. The format reward has no measurable effect because every variant already parses on all localization examples.
  • Precision and recall both improve under GRPO: Precision rises from 0.415 to 0.466 and recall from 0.545 to 0.552 on the primary suite.
  • Grid resolution trade-off: Trained pixel-IoU is similar at N=12 and N=16 (0.235 and 0.231) and falls to 0.192 at N=32. Model-free pixel-IoU is higher at N=16 than N=12 (0.615 versus 0.535), while the 95th-percentile target is shorter at N=16 than N=32 (471 versus 1,926 tokens). N=16 was chosen to balance spatial detail and output length.
  • Boxes win on F1 in one ablation: Bounding boxes achieve higher F1 in the matched-budget localization-only ablation (0.463 versus 0.397), but sparse cells were chosen for their simple text format and direct cell-level reward computation.
  • Parse failures matter for comparisons: For the 900-example score comparison, unparseable scores exclude 154 examples for GPT-5.6-terra and at most 2 for each other API model; assigning the scale midpoint to these failures reduces GPT-5.6-terra's aggregate SRCC to 0.357.

Methodology in Plain English

The authors start from Qwen3-VL-8B and teach it to answer evaluation requests in a structured text format. The generated image is divided into an N × N grid, with N set to 16 by default. Instead of producing masks or heatmaps, the model writes out defect cells by row using 1-based indices, for example r5: 8,9!, where the exclamation mark indicates a cell with high annotation coverage. Separate headers distinguish two defect categories: visual artifacts and semantic misalignments. Empty grids are written as none.

Spatial annotations from the source datasets — masks and heatmaps — are converted into this grid form by resizing and thresholding (masks at 0.5, heatmap-derived coverage marks at ≥ 0.9). Scores from different sources are normalized onto a shared 0–10 scale.

Training happens in two stages. First, supervised fine-tuning with autoregressive cross-entropy on roughly 38K examples, using LoRA (rank 32, scaling 16, dropout 0.05), a learning rate of 2 × 10⁻⁴ with cosine decay and 3% warmup, effective batch size 32, and five epochs. Missing supervision is treated differently from an empty grid: score-only examples simply omit the grid field and are never trained to output none. Clean examples carrying localization supervision are capped at 10% of the corpus.

Second, GRPO post-training from the epoch-4 checkpoint, with the adapter merged into the backbone for full-parameter training. For each input the model samples four responses, which are parsed into scores and defect cells and scored against the available ground truth. The total reward combines a cell-level Dice term (weight 1), a score term (weight 0.3), and a format term (weight 0.1). The Dice reward is coverage-weighted, giving weight 2 to cells marked as high-coverage and 1 otherwise, with β = 1 chosen because it achieved the highest development grid IoU. GRPO settings include 8-bit AdamW at learning rate 10⁻⁵, one epoch, 2,400 stratified training examples, eight gradient-accumulation steps, no KL penalty, and a 768-token completion limit. Inference uses greedy decoding with a 1,024-token output limit.

Finally, a fixed set of rules groups four-connected defect cells, describes up to three components by size with location and coordinate bounds, and labels scores below 5, 5 to below 8, and at least 8 as low, moderate, and high. This parser has no trainable parameters.

Why This Matters

The paper targets a practical gap: generative and editing models fail in localized ways, and scalar scores cannot tell a user where something went wrong. By grounding evaluation in explicit image regions and making the output both machine-parseable and human-readable, VIEScore2 makes automated evaluation auditable rather than opaque. The result that a cell-level Dice reward aligns post-training directly with the model's inference-time output format is a useful methodological pattern for other structured prediction tasks.

Real-world applications:

  • Quality gates for AI image generation and editing products, where localized feedback can guide a user toward the specific region needing a retry.
  • Dataset curation and filtering for training or fine-tuning image models, where region-level defect labels support more selective data cleaning.
  • Editorial, stock-photo, and e-commerce pipelines that need automatic screening of synthetic or retouched imagery for artifacts and prompt violations.
  • Reducing dependence on costly human rating for image-quality evaluation, with the paper noting that manual defect identification is costly, inefficient, and hard to scale.

Industry relevance: The method produces scores, defect grids, and explanations within a single model pass and supports generation and editing tasks with zero, one, or multiple conditioning images, avoiding separate prediction heads per task. That flexibility maps onto the mixed workloads of commercial generative-image platforms, where text-to-image, instruction editing, and reference-based generation all coexist.

Future Directions

  • Backbone and scale generalization: VIEScore2 was validated only with Qwen3-VL-8B, so effectiveness across other VLM backbones and model scales is unknown.
  • Disentangling comparison factors: Comparisons with released evaluators do not separate the effects of architecture, training data, and training procedure.
  • Finer spatial fidelity: The fixed grid can miss small defects and only coarsely approximates irregular boundaries; adaptive or higher-resolution representations are an open problem, given that N=32 reduced trained pixel-IoU to 0.192.
  • Clean-image false alarms and explanation verification: GRPO can increase false alarms on some clean-image subsets, reflecting a trade-off between stronger localization and conservative clean-image recognition, and the parser produces explanations that are faithful to the predictions but do not independently verify their correctness or identify defect causes.

Target Audience

Researchers and engineers working on image quality assessment, multimodal evaluation, or reinforcement-learning post-training for vision-language models; practitioners building automatic QA into generative image pipelines; and readers interested in structured, verifiable output representations that align training objectives with inference-time formats.

Authors’ abstract

Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.

Read the original paper