Skip to content
AI.info

Research

Prominence-Aware Artifact Detection and Dataset for Image Super-Resolution

Overview Research area: Computer vision, specifically single-image super-resolution (SISR) and the detection of visual artifacts produced by generative SR models. Technical level: Intermediate. The pa

Prominence-Aware Artifact Detection and Dataset for Image Super-Resolution
arXiv
2510.16752
Published
2025-10-19
Authors
Ivan Molodetskikh, Kirill Malyshev, Mark Mirgaleev, Nikita Zagainov, Evgeney Bogatyrev, Dmitriy Vatolin

AI summary

Overview

  • Research area: Computer vision, specifically single-image super-resolution (SISR) and the detection of visual artifacts produced by generative SR models.
  • Technical level: Intermediate. The paper assumes familiarity with super-resolution, GANs, diffusion models, and full-reference image-quality metrics, but its central idea — that artifacts should be graded by how noticeable they are — is accessible without deep technical background.
  • Scope: The paper introduces a crowdsourced dataset of SISR artifacts annotated with perceptual prominence scores, a lightweight regressor that predicts spatial prominence heatmaps, and evaluations showing this method outperforms existing artifact detectors and better guides SR fine-tuning.

What This Paper Is About

Generative super-resolution models frequently produce visual artifacts — unnatural patterns, smeared faces, distorted textures — but existing detection methods treat every artifact as an equally important binary defect. In reality, viewers barely notice some artifacts (distortions on water or grass) while others are highly disturbing (distortions on faces or regular structures such as buildings). The authors argue that artifacts should be characterized by their prominence to human observers, and they build both the data and the model needed to do so.

Key Contributions

  1. A new prominence-annotated artifact dataset. The authors present what they describe as the first dataset of its kind: 1302 SISR artifact examples from 11 super-resolution methods, derived from 500-annotated source images (2101 source photos selected from Open Images), each labeled with a crowdsourced prominence score in addition to a binary mask. They also propose a simple mask-postprocessing algorithm to make the tight masks produced by detectors visually judgeable by annotators.

  2. Prominence annotations for the existing DeSRA dataset. They collected crowdsourced prominence scores for all 593 artifact examples in the DeSRA dataset, finding that 48% of them go unnoticed by most viewers.

  3. A perceptual-prominence metric. They train a lightweight regressor that fuses three existing metric features into spatial prominence heatmaps, and report that it outperforms prior artifact detectors under both their own prominence-weighted evaluation and the DeSRA evaluation methodology. They also show that real-time SR output can serve as a pseudo-ground-truth for full-reference metrics when high-resolution ground truth is unavailable.

  4. An analysis of 11 SR methods' artifact proneness. Using their detector alongside existing alternatives, they rank SR methods by how often they generate prominent artifacts, and show that even the newest high-quality methods such as SUPIR are highly susceptible to this problem.

Main Findings

  • Most DeSRA artifacts are barely noticeable. Prominence annotations collected for all 593 DeSRA artifacts showed that 48% go unnoticed by most viewers, and the dataset's distribution centers around moderate prominence. The authors conclude that binary masks are insufficient for accurate SR-artifact evaluation.

  • Their dataset is skewed toward low prominence. The authors state this bias exists because their masks were seeded with results from image-quality metrics poorly adapted to finding artifacts and from detectors that cannot distinguish barely visible from highly visible artifacts. They note this can skew objective evaluation and induce overfitting during training.

  • Mask postprocessing improves annotator judgments. Running the crowdsourced annotation twice on DeSRA — once with postprocessing and once with unmodified masks, using separate participant groups and matching question order — produced a mean artifact prominence of 49.4% with postprocessing versus 47.7% with the original masks.

  • Their method achieves the best average ranking in objective evaluation. On the full proposed and DeSRA datasets, "Ours" at threshold t=0.15 achieved an average PR-AUC ranking of 2.0, ahead of DeSRA (2.3), SSIM (3.7), DISTS (4.2), LDL (4.7), LPIPS (5.2), ssm_jup (6.8), bd_jup (7.2), ERQA (9.8), TOPIQ (10.0), and PaQ-2-PiQ (10.2).

  • The advantage holds on the prominent subset. Restricting to artifacts with prominence above 50% and using the DeSRA comparison methodology (binary IoU, no margin), "Ours" at t=0.15 again ranked best with an average rank of 1.5, ahead of SSIM (2.3), DeSRA (3.3), LPIPS (4.7), LDL (5.0), and DISTS (5.0).

  • Pseudo-ground-truth is viable. The authors report that using lightweight real-time SR methods such as SPAN and RLFN as pseudo-GT for full-reference metrics causes only a small drop in artifact-detection performance compared with using original HR frames, making the approach usable when HR ground truth is unavailable.

  • Artifact representations transfer across SR architectures. In cross-validation where models were trained on two SR families and tested on a held-out third, their approach produced the most consistent results: IoU 0.239 / PR-AUC 0.470 when validating on CNN, 0.231 / 0.419 on Transformer, and 0.331 / 0.482 on Diffusion.

  • Artifact proneness varies widely across SR methods. Crowdsourced evaluation of masks found per SR method showed DRCT (43 masks found, 11.85% mean prominence, 1 confident mask) and HAT-L (53, 13.53%, 1) performing best, while at the other end RealESRGAN (61, 48.42%, 19), SUPIR (70, 45.29%, 20), SwinIR (59, 41.09%, 17), and GFPGAN (45, 32.74%, 11) were most artifact-prone.

  • Their detector finds prominent artifacts efficiently. Among detection methods, their approach found 99 masks with 41.25% mean prominence and 38 confident masks (combined score 15.67%), sharing first place with DISTS in confident masks found (38, at 38.80% mean prominence, combined score 14.75%) while beating it on mean prominence. DeSRA ranked third (110 masks, 32.03% mean prominence, 31 confident, combined 9.93%). LDL at t=0.005 found highly visible artifacts (77.11% mean prominence) but only 12 masks and 11 confident ones, giving a combined score of 8.48%.

  • Their method best guides SR fine-tuning. Fine-tuning RealESRGAN, LDL, and SwinIR against artificial GT built from each detector's masks, their method achieved the best average rank (1.6), ahead of LPIPS (2.2), DeSRA (2.9), DISTS (3.3), LDL (4.9), and ERQA (6.0).

  • Annotation count was chosen empirically. A bootstrap analysis on 11 images annotated by 264 people each (1000 bootstrap repetitions per assessor count k from 1 to 100) showed confidence intervals of roughly ±10% at 100 assessors, while 1–5 assessors frequently spanned the whole 0–100% range. They chose 30 assessors as a compromise, giving roughly ±20% confidence.

Methodology in Plain English

The authors start from a simple premise: a detector that flags every artifact equally will waste effort on invisible defects and miss disturbing ones. To fix this, they first need data that records how noticeable artifacts actually are.

They selected 2101 photos from Open Images, each 768×1024 pixels, downsampled them 4× with bicubic interpolation, and upsampled the results using 11 SR methods to produce 23,111 images to search for artifacts. Initial artifact masks came from manual annotation plus outputs of existing quality metrics (SSIM, DISTS, LPIPS) and detectors (LDL, DeSRA, and in-progress versions of their own method), with the 100 strongest distortions selected per algorithm and images without artifacts manually discarded. The surviving images went to crowdsourced prominence annotation, yielding 697 examples, and a later evaluation step added 605 more.

For the crowd work, they used Toloka.ai. Participants saw image pairs labeled "Original" and "Upscaled" with the artifact region highlighted, and were asked whether the highlighted region contained a distorted object or texture. Every image was ranked by 30 participants, and prominence was computed as the proportion of votes confirming the artifact. Participants had to pass four training questions and four hidden test questions, and every group of 20 questions contained 4 random control questions; responses from anyone failing a control question were discarded. In total, 596 participants completed the annotation.

Because detection methods output tight masks that are hard for humans to judge, they applied morphological postprocessing: opening with a 25×25 square kernel, dilating with a 64×64 circular kernel, and closing with a 25×25 square kernel.

For the detection model itself, they combine three features. The first is DISTS, computed block-wise in 16×16-pixel blocks. The second, called ssm_jup, adapts a small-color-artifact detector (itself based on LDL) to use all RGB channels rather than only chrominance, targeting texture distortions. The third, bd_jup, combines block-wise LPIPS (32×32 blocks, stride 16) with ERQA (8×8 blocks, no overlap), weighting LPIPS 3:2 against ERQA and inverting ERQA so higher means more distortion. These features are fused by a shallow multilayer perceptron with three fully connected layers (3-128-128-1) with ReLU activations, applied per pixel.

Training used Adam on a subset of 374 artifact examples. The loss has two L2 terms: one matching the mean predicted prominence inside the binary mask to the ground-truth prominence, and one pushing the mean outside the mask to zero. Training converges in around 10–15 epochs, with one epoch taking about 13 seconds on an Nvidia RTX 3090.

To make full-reference metrics usable without HR ground truth, they upscale the low-resolution input with a lightweight artifact-resistant method (RLFN or SPAN) to create a pseudo-GT, then compute metrics between that pseudo-GT and the SR output. RLFN pseudo-GT was adopted for all experiments beyond the initial comparison.

For objective evaluation, they binarize detector outputs and compute precision and recall weighted by prominence, using a margin κ = 0.3, so that an artifact confirmed by more than 30% of viewers counts as worth detecting. They also reproduce the DeSRA evaluation methodology, selecting binarization thresholds by maximizing the Precision × Recall product on the prominent DeSRA subset.

Why This Matters

The work reframes super-resolution artifact detection from a binary segmentation problem into a perceptual-ranking problem. That shift matters because the metrics the field currently optimizes against reward finding any distortion at all, which encourages detectors to flag invisible defects while missing the ones that actually ruin a viewer's experience.

Real-world applications:

  • Consumer photo and video upscaling. Running an SR model on user photos or video can introduce distortions on faces and structured objects; prominence-aware detection lets a system suppress only the artifacts users would notice.
  • Video streaming and restoration pipelines. The paper notes that using a lightweight real-time SR method as pseudo-GT enables the approach without access to high-resolution reference frames, which matters for deployed upscaling services.
  • SR model development and benchmarking. The prominence scores provide a finer-grained target for training and comparing SR models than binary mask IoU, and the paper demonstrates use for fine-tuning existing models to reduce artifact proneness.
  • Other restoration tasks. An evaluation on JPEG AI artifacts (Appendix H) suggests the prominence-modeling idea extends beyond super-resolution to learned compression and other restoration problems.

Industry relevance: Companies deploying SR in photo editors, streaming platforms, game upscaling, and content-restoration tools need quality metrics that align with what viewers actually perceive. The reported finding that nearly half of the artifacts in an existing lab-annotated dataset are invisible to most viewers directly challenges how the field measures artifact-detection progress.

Future Directions

  • Extending prominence modeling to video super-resolution, where temporal artifacts such as flickering add a dimension that still-image methods do not address.
  • Handling semantic artifacts. The authors note that newer, higher-capacity models such as SUPIR are starting to produce semantic artifacts like object replacement rather than simple texture distortions, a shift they say warrants further investigation.
  • Improving artifact mask precision. They acknowledge that their masks are approximate, since delineating exact artifact boundaries is ambiguous even for human annotators.
  • Reducing pseudo-GT false positives. They note the pseudo-GT approach can produce false positives when the lightweight SR model fails to reconstruct fine textures.
  • Guiding SR research toward structured regions where artifacts are most visible, using prominence-aware metrics.

Target Audience

Researchers and engineers working on image and video super-resolution, perceptual image-quality assessment, or artifact detection and removal will gain the most from this paper. It is also relevant to practitioners building SR-based products who need evaluation methods that reflect human perception, and to dataset curators interested in crowdsourced annotation methodology — the paper details its training/test/control question design and its bootstrap analysis for choosing annotator counts. The explicit statement that the work does not attempt to prevent SR methods from generating natural-looking but incorrect details makes it relevant to those considering downstream applications like face or license plate recognition, which the authors place outside their scope.

Authors’ abstract

Generative single-image super-resolution (SISR) is advancing rapidly, yet even state-of-the-art models produce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects vary widely in perceptual impact--some are barely noticeable, while others are highly disturbing--yet existing detection methods treat them equally. We propose characterizing artifacts by their prominence to human observers rather than as uniform binary defects. We present a novel dataset of 1302 artifact examples from 11 SISR methods annotated with crowdsourced prominence scores, and provide prominence annotations for 593 existing artifacts from the DeSRA dataset, revealing that 48% of them go unnoticed by most viewers. Building on this data, we train a lightweight regressor that produces spatial prominence heatmaps. We demonstrate that our method outperforms existing detectors and effectively guides SR model fine-tuning for artifact suppression. Our dataset and code are available at https://tinyurl.com/2u9zxtyh.

Read the original paper