Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing
Organizations: Mohamed bin Zayed University of Artificial Intelligence · Institute of Foundation Models · United Arab Emirates University
Abstract
Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.
Figures & tables
| LoomVideo | FLUX.2-klein | FLUX.2-klein-base | VACE | Qwen-Image-Edit | OmniGen2 | |
| 0.22 | 0.35 | 0.22 | 1.79 | 0.30 | 0.24 | |
| 0.13 | 0.10 | 0.15 | 0.69 | 0.35 | 0.28 |
| Head swap | Face swap | Virtual try-on | Background replacement | ||||||||||
| Approach | Repl. | ID | Kpt. | LPIPS | Repl. | ID | Retr. | Kpt. | LPIPS | DINO | LPIPS | DINO | LPIPS |
| JoyAI | 68 99% +31 | .507 .644 +.137 | .043 .085 +.042 | .218 .215 .003 | 60 97% +37 | .282 .591 +.309 | 39 88% +49 | .057 .105 +.048 | .086 .108 +.022 | .283 .298 +.015 | .099 .101 +.002 | .242 .286 +.044 | .063 .065 +.002 |
| LoomVideo | 79 92% +13 | .608 .665 +.057 | .107 .098 .009 | .281 .275 .006 | 71 84% +13 | .449 .466 +.017 | 63 75% +12 | .083 .086 +.003 | .180 .178 .002 | .174 .183 +.009 | .105 .105 .000 | .135 .258 +.123 | .051 .058 +.007 |
| Klein | 91 100% +9 | .726 .790 +.064 | .097 .153 +.056 | .222 .217 .005 | 98 100% +2 | .704 .820 +.116 | 90 96% +6 | .065 .082 +.017 | .064 .076 +.012 | .363 .372 +.009 | .093 .092 .001 | .629 .630 +.001 | .038 .037 .001 |
| Klein-base | 77 100% +23 | .588 .799 +.211 | .169 .236 +.067 | .326 .261 .065 | 95 100% +5 | .557 .831 +.274 | 74 96% +22 | .129 .223 +.094 | .173 .179 +.006 | .409 .465 +.056 | .150 .160 +.010 | .632 .651 +.019 | .162 .136 .026 |
| VACE | — | .405 .729 +.324 | .198 .200 +.002 | .228 .214 .014 | — | .541 .818 +.277 | 87 98% +11 | .187 .193 +.006 | .018 .019 +.001 | .171 .226 +.055 | .056 .056 .000 | .008 .014 +.006 | .008 .008 .000 |
| Head swap | Face swap | Virtual try-on | Background replacement | |||||||||
| Approach | Ref. | Pose | Qual. | Ref. | Pose | Qual. | Ref. | Pose | Qual. | Ref. | Pose | Qual. |
| JoyAI-Video-Edit | 28 /0/ 0 | 3 /25/ 0 | 5 /23/ 0 | 50 /7/ 0 | 3 /50/ 4 | 5 /50/ 2 | 37 /5/ 12 | 0 /52/ 2 | 5 /46/ 3 | 19 /25/ 6 | 0 /50/ 0 | 1 /46/ 3 |
| LoomVideo | 7 /18/ 2 | 12 /14/ 2 | 8 /14/ 6 | 18 /31/ 3 | 3 /44/ 5 | 1 /33/ 18 | 36 /15/ 0 | 0 /51/ 0 | 3 /48/ 0 | 39 /12/ 0 | 0 /50/ 1 | 1 /49/ 1 |
| FLUX.2-klein | 20 /5/ 0 | 0 /25/ 0 | 7 /18/ 0 | 33 /16/ 2 | 0 /48/ 3 | 1 /48/ 2 | 1 /11/ 1 | 1 /12/ 0 | 0 /12/ 1 | 0 /10/ 0 | 0 /10/ 0 | 0 /10/ 0 |
| FLUX.2-klein-base | 24 /1/ 0 | 0 /25/ 0 | 3 /22/ 0 | 45 /8/ 0 | 2 /45/ 6 | 7 /44/ 2 | 20 /16/ 18 | 0 /53/ 1 | 1 /51/ 2 | 15 /40/ 1 | 0 /56/ 0 | 10 /45/ 1 |
| VACE | 24 /0/ 1 | 0 /24/ 1 | 6 /15/ 4 | 35 /19/ 1 | 5 /48/ 2 | 2 /50/ 3 | 57 /0/ 0 | 5 /52/ 0 | 8 /49/ 0 | — | — | — |
| RefGAP | best admissible edit-side constant | ||||||
| Approach | ID | Kpt. | LPIPS | ID | Kpt. | LPIPS | |
| JoyAI-Video-Edit | .644 | .085 | .215 | 1.5 | .644 | .075 | .207 |
| LoomVideo | .665 | .098 | .275 | 2.5 | .671 | .098 | .271 |
| FLUX.2-klein | .790 | .153 | .217 | 2.0 | .794 | .171 | .229 |
| FLUX.2-klein-base | .799 | .236 | .261 | 2.0 | .801 | .236 | .260 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Denoising (s/sample) | End-to-end (s/sample) | Peak memory | |||||||
| Approach | Steps | Base | RefGAP | Base | RefGAP | Base (GB) | (MiB) | ||
| VACE | 30 | 158.2 | 159.0 | 159.4 | 160.0 | 26.30 | |||
| LoomVideo | 50 | 106.2 | 108.9 | 112.0 | 116.6 | 43.95 | |||
| Qwen-Image-Edit | 20 | 62.48 | 66.00 | 64.0 | 67.4 | 57.95 | |||
| OmniGen2 | 50 | 21.68 | 23.39 | 22.4 | 24.2 | 15.36 | |||
| FLUX.2-klein-base | 50 | 16.35 | 20.34 | 17.0 | 21.2 | 15.55 | |||
| Approach | |||||||
| JoyAI-Video-Edit | n/a | .0615 | — | .0555 | — | 2.30 | 2.96 |
| LoomVideo | .031 | .0069 | 0.22 | .0040 | 0.13 | 2.32 | 5.91 |
| FLUX.2-klein | .286 | .0988 | 0.35 | .0290 | 0.10 | 4.48 | 2.68 |
| FLUX.2-klein-base | .286 | .0635 | 0.22 | .0433 | 0.15 | 1.88 | 3.24 |
| VACE | .033 | .0590 | 1.79 | .0228 | 0.69 | 5.99 | 3.21 |
| Qwen-Image-Edit | .332 | .0990 | 0.30 | .1157 | 0.35 | 1.11 | 2.47 |
| ID | LPIPS ne | ID | LPIPS ne | ID | LPIPS ne | |
| JoyAI-Video-Edit ( ); base ID , LPIPS ne | ||||||
| all layers | ||||||
| .689 | .0615 | .683 | .0625 | .682 | .0618 | |
| .705 | .0635 | .705 | .0644 | .706 | .0641 | |
| .548 | .0913 × | .543 | .0914 × | .549 | .0917 × | |
| JoyAI-Video-Edit | LoomVideo | FLUX.2-klein | FLUX.2-klein-base | |||||
| all | second half | second half | all | |||||
| ID | LPIPS ne | ID | LPIPS ne | ID | LPIPS ne | ID | LPIPS ne | |
| base | .557 | .0663 | .653 | .0700 | .807 | .0435 | .590 | .1206 |
| 0.5 | .647 | .0641 | .661 | .0681 | .851 | .0377 | .806 | .0667 |
| 1.0 | .683 | .0625 | .689 | .0676 | .883 | .0394 | .861 | .0659 |
| 1.5 | .703 | .0616 | .717 | .0669 | .888 | .0410 | .880 | .0667 |
| ID (AdaFace) | Kpt. | LPIPS | |
| DirectSwap |
| JoyAI-Video-Edit | LoomVideo | FLUX.2-klein | FLUX.2-klein-base | VACE | Qwen-Image-Edit | OmniGen2 | DirectSwap | |
| base | 0.809 | 0.816 | 0.861 | 0.757 | 0.786 | 0.734 | 0.581 | 0.894 |
| RefGAP | 0.876 | 0.861 | 0.887 | 0.875 | 0.888 | 0.802 | 0.851 | 0.894 |
| win % | 73 | 73 | 61 | 76 | 96 | 63 | 96 | 49 |
| ID-Retr. (%) | ID-sim | Kpt. | ||||
| input | base | RefGAP | base | RefGAP | base | RefGAP |
| full frame ( ) | 25.0 | 50.0 | .109 | .240 | .085 | .108 |
| face-centered crop ( ) | 40.0 | 80.0 | .248 | .434 | .058 | .074 |
| Approach | DINO | LPIPS |
| FLUX.2-klein | 0.382 0.391 | 0.0302 0.0318 |
| FLUX.2-klein-base | 0.407 0.426 | 0.0415 0.0417 |
| Qwen-Image-Edit | 0.216 0.225 | 0.0158 0.0160 |
| OmniGen2 | 0.397 0.443 | 0.1339 0.1395 |
| Task | Metric | JoyAI | LoomVideo | Klein | Klein-base | VACE | Qwen | OmniGen2 |
| Head swap | Repl. (pp) | +30.9 [+21.5, +40.1] | +12.2 [+6.6, +17.6] | +8.6 [+5.1, +12.3] | +22.8 [+17.5, +27.4] | — | +16.1 [+10.1, +24.1] | +65.1 [+56.7, +71.4] |
| ID | +.137 [+.111, +.169] | +.058 [+.045, +.069] | +.064 [+.049, +.088] | +.211 [+.176, +.247] | +.323 [+.290, +.362] | +.114 [+.071, +.167] | +.486 [+.449, +.525] | |
| Kpt. | +.042 [+.027, +.059] | .009 [ .025, +.002] | +.055 [+.046, +.065] | +.068 [+.048, +.096] | +.001 [ .009, +.012] | +.000 [ .017, +.015] | +.002 [ .019, +.021] | |
| LPIPS | .003 [ .010, +.006] | .005 [ .013, +.000] | .004 [ .011, +.005] | .066 [ .093, .037] | .013 [ .016, .011] | .212 [ .246, .171] | .099 [ .117, .079] | |
| Face swap | Repl. (pp) | +36.8 [+31.4, +42.6] | +13.0 [+9.4, +17.1] | +2.2 [+0.8, +4.1] | +5.1 [+2.9, +7.8] | — | +16.7 [+11.7, +21.8] | +20.7 [+16.7, +24.9] |
| ID | +.310 [+.293, +.327] | +.017 [+.003, +.034] | +.116 [+.102, +.131] | +.274 [+.244, +.305] | +.277 [+.262, +.294] | +.169 [+.137, +.202] | +.498 [+.474, +.521] |
| RefGAP | base | both succeed | both fail | tie % | |
| Reference fidelity | 78 (27.9%) | 10 (3.6%) | 144 | 48 | 68.6 |
| Pose / expression | 20 (7.1%) | 10 (3.6%) | 178 | 72 | 89.3 |
| Visual quality | 0 (0.0%) | 22 (7.9%) | 246 | 12 | 92.1 |
| Task | Question | RefGAP | both succeed | both fail | base | total |
| Head swap | closer to reference | 157 | 23 | 1 | 3 | 184 |
| pose / expression | 27 | 109 | 46 | 3 | 185 | |
| visual quality | 37 | 131 | 7 | 10 | 185 | |
| Face swap | closer to reference | 278 | 89 | 7 | 8 | 382 |
| pose / expression | 19 | 185 | 138 | 40 | 382 | |
| visual quality | 32 | 312 | 8 | 30 | 382 |
| Task | Ref. | Pose/expr. | Quality | Ties |
| Head swap | 93.5 [84.7,100] | 86.0 [66.0,100] | 86.2 [71.9,97.9] | 12.6 [3.6,23.4] |
| Face swap | 93.6 [87.2,98.7] | 39.7 [23.0,57.0] | 56.1 [40.0,72.2] | 26.0 [16.9,35.4] |
| Try-on | 85.2 [75.4,93.8] | 71.4 [42.9,100] | 85.0 [70.0,100] | 14.6 [6.8,23.5] |
| Background | 82.2 [68.8,93.8] | 40.0 [0.0,80.0] | 52.2 [30.0,74.4] | 45.7 [32.7,58.6] |
| Approach | kpt | ID | lowest ID quartile | highest ID quartile | ( ) |
| JoyAI-Video-Edit | (44) | ||||
| LoomVideo | (350) | ||||
| FLUX.2-klein | (348) | ||||
| FLUX.2-klein-base | (230) | ||||
| VACE | (13) | ||||
| Qwen-Image-Edit | (338) |
| HeadSwapBench | ID | kpt | LPIPS |
| FLUX.2-klein (4-step schedule) | |||
| base | 0.726 | 0.097 | 0.222 |
| RefGAP ( , all steps) | 0.790 | 0.153 | 0.217 |
| RefGAP, skip step 1 | 0.753 | 0.101 | 0.221 |
| JoyAI-Video-Edit (2-step schedule) | |||
| base | 0.507 | 0.043 | 0.218 |
| diagnostic clips ( ) | all clips ( ) | ||||
| setting | layers | edit-side strength | DINO | subject LPIPS | DINO |
| base | — | — | .473 | .0384 | .299 |
| RefGAP (shared setting) | all | online, | .112 | .0242 | .042 |
| RefGAP, second half | – | online, | .106 | .0127 | — |
| fixed edit-side bias | all | constant, | .040 | .0174 | .037 |