cs.CVSep 28, 2026

Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing

Authors: Yanan Wang, Shengcai Liao, Guangyi Liu, Xiaodan Liang

Organizations: Mohamed bin Zayed University of Artificial Intelligence · Institute of Foundation Models · United Arab Emirates University

Abstract

Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Toward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention

    Jun 12, 2026Dinh-Khoi Vo, Nhut-Thanh Le-Hinh, Viet-Tham Huynh +3Diffusion-Based Image EditingOmnidirectional Images

  2. Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

    Sep 17, 2026Yulong Chen, Ziqian Zhang, Haoyu Zhang +4Diffusion-Based Image EditingImage Editing

  3. Masked Generative Transformer Is What You Need for Image Editing

    May 11, 2026Wei Chow, Linfeng Li, Xian Sun +14Diffusion-Based Image EditingImage Editing