Research
AI-Generated Image Detectors Overrely on Global Artifacts: Evidence from Inpainting Exchange
Overview Research area: Computer vision, specifically AI-generated image detection and image forensics for diffusion-based inpainting. Technical level: Intermediate. The paper's central idea is intuit

- arXiv
- 2602.00192
- Published
- 2026-01-30
- Authors
- Elif Nebioglu, Emirhan Bilgiç, Adrian Popescu
AI summary
Overview
Research area: Computer vision, specifically AI-generated image detection and image forensics for diffusion-based inpainting.
Technical level: Intermediate. The paper's central idea is intuitive, but the supporting arguments rely on assumptions about diffusion models, VAE latent compression, KL divergence, and Fourier-domain spectral analysis.
Scope: The paper argues that state-of-the-art AI-image detectors succeed largely because diffusion inpainting leaves a global spectral fingerprint across the whole image, and it introduces a pixel-level "Inpainting Exchange" operation that removes that fingerprint to test whether detectors actually recognize generated content.
What This Paper Is About
Diffusion-based inpainting edits only a local masked region, so a competent detector should be looking at the generated pixels inside that region. Instead, the authors find that detectors are keying on a global artifact: the VAE encoder-decoder at the heart of latent diffusion reconstructs the entire image, subtly altering even the unedited pixels and shaving off high-frequency detail everywhere. The paper's goal is to isolate and measure this shortcut by restoring the original background pixels after inpainting, leaving the synthesized content untouched, and then observing how detectors behave.
Key Contributions
-
INP-X (Inpainting Exchange): An operation that restores the original pixels outside the edited mask while preserving all synthesized content inside it, producing images where any detector signal must come from the generated region itself.
-
Theoretical analysis plus empirical validation: Theorem 3.2 links the observed global artifact to high-frequency attenuation caused by the VAE information bottleneck, and Theorem 3.4 shows the exchange operation sets the background term of the KL divergence between real and manipulated distributions to exactly zero. The theory is checked experimentally using the Stable Diffusion v1.4 VAE in isolation.
-
A 90K-image benchmark: Built by extending Semi-Truths across 4 datasets (CelebA-HQ, CityScapes, OpenImages, SUN-RGBD) and 3 inpainting models (Kandinsky 2.2, OpenJourney, Stable Diffusion v1.4), with matched real / standard-inpainted / exchanged triplets and masks. Each of the three subsets contains 30k samples.
-
A broad evaluation: 11 open-source pretrained detectors and 2 commercial APIs (HiveModeration, Sightengine) are tested on standard inpainting versus INP-X, plus 4 fine-tuned detectors (CLIP ViT B/32, ViT B/16, EfficientNet, ResNet-50) evaluated on both classification and localization.
Main Findings
-
Detectors collapse under INP-X. Across the 11 pretrained detectors, accuracy on standard inpainting ranged from 0.502 (Cross-EfficientViT) to 0.942 (Corvi2023); under INP-X it ranged from 0.501 to 0.604 (DNF being the best at 0.604). Several detectors were already near chance on standard inpainting.
-
Commercial APIs drop to roughly chance. Hive Moderation fell from 0.914 accuracy on standard inpainting to 0.548 on INP-X; Sightengine fell from 0.926 to 0.550. The abstract paraphrases this as a drop from above 91% to around 55%.
-
Frequency-based detectors suffer most. Corvi2023, which relies on spectral analysis, fell from 0.942 to 0.554, and its recall dropped from 0.890 to 0.114, supporting the claim that it exploits the global spectral artifact.
-
The degradation is mostly a recall failure. The paper attributes most of the accuracy loss to low recall rather than low precision, meaning detectors stop flagging manipulated images rather than mislabeling clean ones.
-
The VAE is the source of the artifact. Isolating the Stable Diffusion v1.4 VAE and encoding/decoding real images with no diffusion reproduces the same spatial structure seen in inpainting differences. At the image level, VAE reconstruction loss correlates with high-frequency content at Pearson r = 0.941, and with inpainting loss at r = 0.600. Pixel-level correlations are weaker but consistent (VAE↔Inpaint r = 0.45 ± 0.23, VAE↔HighFreq r = 0.52 ± 0.08, Inpaint↔HighFreq r = 0.33 ± 0.15).
-
Fine-tuning helps, but INP-X remains hard. The best INP-X-trained detector reached 0.753 accuracy (CLIP ViT B/32). Models trained on standard inpainting then tested on INP-X dropped as low as 0.557 (ResNet-50) and 0.567 (EfficientNet).
-
Transfer is asymmetric. Detectors trained on INP-X transfer better to standard inpainting (ResNet-50 reaches 0.745) than the reverse, suggesting local-artifact training yields more general features than global-artifact training.
-
Localization improves under INP-X training. Training on exchanged images improved mIoU and mAP relative to INP-trained baselines across all four architectures; the authors describe the gains as modest but significant (p ≤ 0.05, t-test), and larger for CNN backbones than for transformers.
-
INP-X is stronger than standard corruptions. Gaussian blur (σ = 3, μ = 0), a localized "Gaussian Light Spot" attack (radius r = 120, intensity multiplier A = 1.5), and JPEG compression at quality 80 degraded some detectors but generally far less than INP-X. Commercial APIs stayed above 72% accuracy (Sightengine) and above 89% (Hive Moderation) under those corruptions while sitting near 55% under INP-X.
-
The spectral fingerprint is suppressed. Using a pixel-wise cross-difference high-pass filter followed by 2D FFT on images resized to 512 × 512, spectral MSE (×1000) fell from 15.3093 to 12.8710 (CelebA-HQ), 11.4707 to 5.8937 (CityScapes), 17.1325 to 1.9152 (OpenImages), and 23.4154 to 1.9436 (SUN-RGBD) — the paper highlights an 11× improvement on SUN-RGBD.
-
The effect is not simply a matter of mask size. CelebA-HQ has a larger mean mask ratio (μ_m ≈ 0.10) than SUN-RGBD (μ_m ≈ 0.06) yet shows a narrower spectral gap. Detection accuracy generally rises with mask size, but INP-X degrades performance at every size.
Methodology in Plain English
The researchers start from a suspicion: a diffusion inpainter bakes the whole image through a VAE encoder and decoder, so even the pixels it was told not to touch come back slightly changed. To test whether detectors rely on that change rather than the fake content, they build INP-X. Take a real image, run standard inpainting on it to produce a manipulated version, then paste the original pixels back everywhere outside the mask. The fake region stays exactly as the model generated it, but the background is byte-for-byte real again.
They then construct a 90K benchmark with three matched versions of each scene — real, standard-inpainted, exchanged — across four source datasets and three inpainting models, each paired with its mask. Detectors are run on both the standard-inpainted and the exchanged versions to see whether performance holds up when the background artifact is gone.
To justify the mechanism, they first prove that a VAE trained with an MSE-dominated objective cannot encode high-frequency sensor noise (it falls into the irreducible residual of the posterior variance), so reconstructed images lose high-frequency power. They then show via the chain rule for KL divergence that standard inpainting introduces divergence in the background term, whereas exchange forces that term to zero by construction. As a check, they encode and decode real images with the Stable Diffusion v1.4 VAE alone, with no diffusion step, and correlate that pure reconstruction error against high-frequency content and against the inpainting difference.
Finally, they fine-tune four detector backbones on INP and INP-X data separately and test them on each other's distributions, measuring both binary classification and localization against ground-truth masks (saliency maps thresholded at 0.5, masks resized to 224 × 224). They also run corruption robustness tests and FFT-based spectral comparisons, and vary mask sizes to rule out simple explanations.
Why This Matters
The paper reframes a high-accuracy result as a measurement artifact. If detectors are reading a pipeline fingerprint rather than the manipulated content, then their reported accuracy says little about whether they can catch real manipulations, and their reliability will change as generation pipelines change.
Impact on research: It provides a controlled intervention that exposes shortcut learning in the inpainting setting and supplies both a benchmark and a diagnostic tool, arguing that evaluation should include realistic post-edits rather than only "clean" generated images.
Real-world applications:
- Misinformation and content authenticity review, where a detector that cannot separate generated content from a global artifact may misjudge partially edited images.
- Newsroom and fact-checking workflows, which increasingly rely on detector scores as one signal among several.
- Platform content moderation, where near-chance behavior under realistic edits could mean manipulated images pass review.
- Forensic and legal image authentication, where the paper's argument implies that sensor-noise-based authenticity checks (such as PRNU) behave differently on INP-X than on standard inpainting.
Industry relevance: The two commercial APIs tested are named systems that dropped from above 91% to roughly 55% accuracy under INP-X, which is directly relevant to vendors selling detection-as-a-service. The paper also releases its dataset and code publicly at github.com/emirhanbilgic/INP-X.
Future Directions
-
Better detection training. The authors suggest that detectors trained on INP-X-style data learn local content features and propose content-aware, localization-based detection as the path forward.
-
Frequency-preserving VAEs. Training or decoding strategies that explicitly enforce spectral consistency could remove the artifact at its source rather than only diagnosing it.
-
Architectures beyond VAEs. The paper states as a limitation that its dataset focuses on VAE-based architectures; component-based or pixel-space models, described as currently less efficient and significantly less used, are flagged as needing deeper analysis.
-
Open problem of INP-X detection itself. Even detectors trained specifically on exchanged images top out at 0.753 accuracy, so detecting this kind of edit remains unsolved.
-
Boundary and blending quality. The authors note that subtle boundary effects may occur depending on mask and blending, that alpha blending mitigates them, and that they do not claim perceptual quality — leaving room for better blending approaches such as Poisson editing.
Target Audience
Researchers working on AI-generated image detection, image forensics, and generative model evaluation; practitioners evaluating or deploying commercial detection APIs; and anyone building benchmarks for media authenticity. It is most useful to readers who already understand how latent diffusion models work and are comfortable with the idea that a classifier might be solving an easier problem than the one it was built for.
Authors’ abstract
Modern deep learning-based inpainting enables realistic local image manipulation, raising critical challenges for reliable detection. However, we observe that current detectors primarily rely on global artifacts that appear as inpainting side effects, rather than on locally synthesized content. We show that this behavior occurs because VAE-based reconstruction induces a subtle but pervasive spectral shift across the entire image, including unedited regions. To isolate this effect, we introduce Inpainting Exchange (INP-X), an operation that restores original pixels outside the edited region while preserving all synthesized content. We create a 90K test dataset including real, inpainted, and exchanged images to evaluate this phenomenon. Under this intervention, pretrained state-of-the-art detectors, including commercial ones, exhibit a dramatic drop in accuracy (e.g., from 91\% to 55\%), frequently approaching chance level. We provide a theoretical analysis linking this behavior to high-frequency attenuation caused by VAE information bottlenecks. Our findings highlight the need for content-aware detection. Indeed, training on our dataset yields better generalization and localization than standard inpainting. Our dataset and code are publicly available at https://github.com/emirhanbilgic/INP-X.