cs.CVOct 6, 2026

Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization

Authors: Xuekang Zhu, Kaiwen Feng, Ruifeng Wang, Xiwen Wang, Xiaochen Ma, Bo Du, Changjiang Jiang, Chenfan Qu, +5 more

Organizations: Sichuan University · Ant Group · The Hong Kong University of Science and Technology · Wuhan University · South China University of Technology · University of Southern California · Xiamen University of Technology

Abstract

Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask yy for a given image xx. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, P(y∣x)=∫P(y∣z) P(z∣x) dzP(y|x)=\int P(y|z)\,P(z|x)\,dz, where zz denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling zz in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate P(z∣x)P(z|x) and P(y∣z)P(y|z) via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at https://github.com/venus-guangjian/PAL

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Courtroom Trial of Pixels: Robust Image Manipulation Localization via Adversarial Evidence and Reinforcement Learning Judgment

    Apr 16, 2026Songlin Li, Zhiqing Guo, Dan Ma +2PixelsAdversarial Robustness

  2. SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

    May 27, 2026Peiyu Zhuang, Jianquan Yang, Haodong Li +6Neural Mask Estimation