Skip to content
AI.info

Research

Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization

Overview Research area: Computer vision, specifically image forensics and Image Manipulation Localization (IML) — finding which pixels in an image have been altered. Technical level: Advanced. The pap

Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization
arXiv
2610.07916
Published
2026-10-06
Authors
Xuekang Zhu, Kaiwen Feng, Ruifeng Wang, Xiwen Wang, Xiaochen Ma, Bo Du, Changjiang Jiang, Chenfan Qu, Songyu Ye, Xia Du, Wentao Feng, Jian Liu, Ji-Zhe Zhou

AI summary

Overview

Research area: Computer vision, specifically image forensics and Image Manipulation Localization (IML) — finding which pixels in an image have been altered.

Technical level: Advanced. The paper reformulates IML as a latent-variable inference problem and derives a Bayes-optimal justification for its training objective.

Scope: This paper argues that manipulation "artifacts" should be modeled as an explicit latent variable rather than learned implicitly, and proposes a two-stage training framework (PAL + SL) plus a new 45K-group dataset (EditGroup-45K) to do so.

What This Paper Is About

Most image manipulation localization models are trained end-to-end to map an image directly to a tampered-region mask, assuming the network will pick up manipulation artifacts on its own. The authors argue that artifacts — the subtle, non-semantic traces left by editing — are invisible to the eye and impossible to annotate directly, so they behave like a hidden (latent) variable that current models only capture implicitly, leaving them prone to learning semantic shortcuts instead. Their goal is to model that latent artifact variable explicitly by comparing related and unrelated image pairs, then reuse what was learned as a better starting point for standard localization models.

Key Contributions

  1. Explicit artifacts modeling paradigm. The authors identify artifacts as a latent variable and reformulate IML as the probabilistic decomposition P(y|x) = ∫ P(y|z)P(z|x) dz, where z denotes the artifacts. They argue the failure mode of existing methods is their implicit artifacts modeling strategy, and that z should instead be a direct optimization target.

  2. Two-stage disentangled learning framework. They propose a paradigm consisting of Pairwise Artifacts Learning (PAL) and Standard Localization (SL). PAL learns the artifacts posterior P(z|x) by contrasting edit-related and non-edit-related image pairs; SL then adapts the resulting encoder to the pixel-level mask objective P(y|z).

  3. EditGroup-45K dataset. A source-anchored dataset organized into edit groups, comprising 45,546 aligned edit groups, 52,074 authentic images, and 232,520 manipulated variants (described in the introduction as roughly 284K images, and elsewhere as approximately 285K). It contains 140,137 manual edits (60.27%) and 92,383 AIGC edits (39.73%), spanning copy-move (70,958), splicing (60,603), inpainting (50,911), and full-image generation (50,048).

  4. Performance boost and validation. They show PAL initialization improves localization across many backbones and existing baselines, and provide quantitative boundary-activation analysis and qualitative visualization to argue that PAL genuinely captures artifacts via feature disentanglement.

Main Findings

  • Pairwise learning beats absolute supervision on generalization. On the ablation benchmark, PAL reaches the highest All-Avg score of 0.6266 with Cross-Avg 0.5868 and In-Avg 0.8256. Absolute binary supervision (Baseline-ABS) scores higher in-domain (In-Avg 0.8457) but its Cross-Avg (0.5292) falls slightly below the plain Baseline (0.5320), which the authors attribute to overfitting to the source domain.

  • Explicit feature difference is the single most important design choice for cross-domain transfer. Removing it ("w/o Diff") drops Cross-Avg from 0.5868 to 0.5333, the largest cross-domain decline among the representation-design variants.

  • Both pooling strategies matter, but for different reasons. Removing average pooling ("w/o AvgPool") produces the lowest In-Avg of 0.7814, indicating global context aggregation is needed for source-domain stability. Removing max pooling ("w/o MaxPool") keeps In-Avg high at 0.8368 but lowers Cross-Avg to 0.5676, suggesting peak local inconsistencies drive cross-domain robustness.

  • Self-augmented negatives trade generalization for in-domain scores. Removing them ("w/o SA") spikes In-Avg to the highest value of 0.8521 — above PAL itself — while Cross-Avg drops to 0.5492, which the authors interpret as the model learning trivial source-domain shortcuts.

  • In-group Edited–Edited negatives are the tightest constraint. Removing them ("w/o IG-EE") causes the most severe generalization decline, giving the lowest Cross-Avg of 0.5386.

  • PAL amplifies boundary activations by roughly 5x to 7x. Measuring feature activations along boundaries extracted via 7x7 morphological erosion, the ImageNet-pretrained ConvNeXt-Small backbone versus its PAL-optimized counterpart shows ratios of 6.32 (CASIAv1, 0.48 to 3.03), 5.00 (Columbia, 0.47 to 2.36), 7.40 (Coverage, 0.46 to 3.38), 5.62 (NIST16, 0.58 to 3.24), 6.64 (AutoSplice, 0.41 to 2.74), and 6.22 (CocoGlide, 0.46 to 2.83).

  • PAL transfers across architectures regardless of inductive bias. Relative gains ranged from +6.27% (Swin-Small, 0.5440 to 0.5781) to +10.68% (ConvNeXt-Small, 0.5661 to 0.6266), with ResNet-101 at +7.12% (0.4878 to 0.5226) and SegFormer-B3 at +8.24% (0.5496 to 0.5949). Foundation-style backbones also improved: DINOv3-Base from 0.6114 to 0.6477 (+5.94% per the text, +5.9 in the table) and SAM-Base from 0.5206 to 0.5937 (+14.0 in the table, +14.04% in the text).

  • Existing IML methods improve when their backbones are swapped for PAL weights. CAT-Net 0.5452 to 0.5812 (+6.60), PSCC 0.5281 to 0.5583 (+5.72), MVSS 0.4946 to 0.5437 (+9.93), Mesorch 0.5509 to 0.5853 (+6.24), and TruFor 0.5623 to 0.6095 (+8.39).

  • BCE is preferred over a contrastive objective. The authors report that BCE converges faster and reaches a higher performance ceiling than a contrastive variant (PAL-Con) that keeps the same pair construction, shared encoder, feature difference, and pooling but replaces the binary classifier. The quantitative comparison is placed in Appendix B, which is truncated in the provided content, so the specific contrastive numbers are not reported here.

Methodology in Plain English

The reasoning starts from a simple observation: an image that has been edited differs from its original both because of semantic changes (the content that was added or removed) and because of artifacts (the forensic traces left behind). If you only ever compare a tampered image against random images, you cannot tell which differences are artifacts and which are just different content.

The authors' fix is to build image groups. Each group is anchored on one or more authentic images, and all other images in the group are edited versions derived from those anchors. This creates a family tree of edits, which in turn lets the researchers construct two kinds of pairs. Positive pairs are an authentic image and its own edited counterpart — their differences include artifacts. Negative pairs are pairs that have no direct editing relationship but are statistically similar in other ways. They designed six negative types to cover this: cross-group original–edited, self-augment, in-group edited–edited, in-group original–original, global edited–edited, and global original–original.

Stage I (PAL) then passes both images in a pair through a shared encoder, takes the absolute difference of their feature maps, aggregates that difference with global average and max pooling into a single relation vector, and feeds it to a lightweight decoder that must output a single number: are these two images edit-related or not? Training uses a pairwise binary cross-entropy loss. The authors show mathematically that the best possible score for this task is the log ratio of the edit-relation distribution to the negative-sampling distribution, and that by choosing negatives whose marginal distribution roughly matches the positives' marginal distribution, the model is forced to rely on the edit-induced traces rather than on what the images happen to depict. That is what they mean by "explicitly disentangling" artifacts.

Stage II (SL) is deliberately ordinary: the encoder trained in Stage I becomes the initialization for a standard localization model, and training proceeds exactly as in the existing IML codebase (Protocol-CAT setting) to produce pixel-level masks.

Why This Matters

Impact on research. The paper challenges the default assumption that a sufficiently large end-to-end model will learn forensic artifacts on its own. It reframes the problem as latent-variable inference, offers a theoretical argument for why pairwise comparison with matched negatives isolates artifacts, and shows the resulting encoder is a general-purpose initialization rather than a bespoke architecture — the gains appear across CNNs, Transformers, and foundation-style backbones, and across methods such as CAT-Net, PSCC, MVSS, Mesorch, and TruFor.

Real-world applications:

  • Journalism and news verification — locating which parts of a submitted photograph were altered before publication.
  • Insurance and legal evidence review — flagging spliced or inpainted regions in submitted claim photos or evidentiary images.
  • Social media and platform content moderation — detecting AI-generated or edited imagery at scale, especially given that 39.73% of EditGroup-45K's edits are AIGC-based.
  • Digital media authentication and provenance tooling — supplying source-anchored relationships that support tracking which image a manipulated variant came from.

Industry relevance. The dataset construction covers both traditional manual editing tools such as Adobe Photoshop (copy-move, splicing, removal) and modern generative editing models including Doubao, Qwen-Edit, Banana, and GPT-Image-1. Because EditGroup-45K supplies source-anchored correspondence at a scale that prior anchored datasets lack — comparable datasets in the paper's comparison table include Fantastic Reality at 16,592 authentic and 19,423 manipulated images, and CIMD at 400 and 400 — it is positioned as training infrastructure for forensic systems that must handle manual and generative manipulation together. Code and dataset are released at https://github.com/venus-guangjian/PAL.

Future Directions

  • Explain the contrastive-versus-BCE gap in more depth. The comparison is deferred to Appendix B with only a qualitative claim (faster convergence, higher ceiling). A fuller characterization of when pairwise BCE fails and when a contrastive objective would be preferable is left open.

  • Extend beyond binary relation scoring. PAL reduces a pair to a single scalar relation score. Whether richer relation outputs — for example, predicting the type of edit (copy-move vs. splicing vs. inpainting vs. full-image generation) — would further sharpen disentanglement is not explored.

  • Scale and diversify the anchor structure. EditGroup-45K's anchors come from M3, with generation via four named models. Whether the marginal-matching construction remains effective as generators change, or whether it must be re-balanced per generator family, is not addressed.

  • Test on manipulation types absent from the group taxonomy. The dataset covers four categories (copy-move, splicing, inpainting, full-image generation). Generalization to manipulation families outside these four is not reported.

Target Audience

Researchers and graduate students working on image forensics, tampered-image detection, and media forensics; practitioners building deepfake and manipulation detection pipelines who want a drop-in encoder initialization; and machine learning researchers interested in latent-variable formulations, feature disentanglement, and pairwise or contrastive training objectives applied to low-level, non-semantic signals. Readers should be comfortable with probabilistic notation such as P(y|x) and likelihood ratios, though the core intuition — compare related pairs against unrelated pairs to isolate what editing leaves behind — is accessible without the derivations.

Authors’ abstract

Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P(y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at https://github.com/venus-guangjian/PAL

Read the original paper