Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (CFG), an editor combines two directions at every denoising step, one that pushes toward the instructed edit and one that pulls back toward the source image, using global weights. We introduce Attention-Scoped Guidance (ASG), a sampler wrapper that makes these weights spatial. It reads a soft support map from the instruction attention that the editor already computes, then weakens text guidance where support is low and strengthens image anchoring where support is high. The wrapper requires no training, no external mask, and no additional network evaluation. On the full MagicBrush and PIE-Bench++ splits, ASG improves preservation-oriented metrics, leading three of four MagicBrush metrics and PIE-Bench++ background PSNR. A dose-matched control that removes the spatial placement loses up to 0.73 CLIP on PIE-Bench++, confirming that the spatial allocation itself carries the gain.
Figures & tables
Figure 1: An illustrative failure from a ten-case probe, motivating Attention-Scoped Guidance. Left: for a background-only MagicBrush instruction, the underlying editor changes both the scene and the subject. Center: the displayed attention maps retain spatial structure through denoising, with nearly unchanged top-10% mass. Right: the dual-CFG combination uses global text and image weights, motivating explicit spatial allocation.
Figure 2: Method overview, left to right. Given a source image and an instruction, the frozen editor evaluates three noise predictions at every denoising step. Instruction cross-attention, averaged over heads and layers and summed over tokens and the first ten steps, yields a soft support M(x) . From the second step on, this support makes the two CFG weights spatial: the text-edit direction is weakened where M is low, and the image anchor is strengthened where M is high; the support freezes after step ten. The reweighted directions are recombined and sampled unchanged, completing the edit within the original sampling trajectory, with no additional network evaluations.
MagicBrush-I
MagicBrush-T
PIE-Bench++
System
CLIP-I ↑
L1 ↓
CLIP-I ↑
L1 ↓
CLIP-W ↑
CLIP-E ↑
PSNR-BG ↑
InstructPix2Pix
0.8509
0.1135
0.8099
0.1491
24.001
23.457
80.656
MagicBrush-IP2P
0.9177
0.0741
0.8849
0.1042
23.612
22.985
81.388
InstructDiffusion
0.9228
0.0706
0.8912
0.0995
23.903
23.155
81.591
MGIE
0.9087
0.0812
0.8745
0.1133
21.357
21.080
79.461
InstructCLIP
0.8794
0.0962
0.8360
0.1291
23.441
22.767
80.985
Table 1: Main full-population comparison against released external editors. MagicBrush-I and MagicBrush-T denote the independent and iterative MagicBrush test protocols. Best values within each metric are bold. The last row reports Attention-Scoped Guidance.
Configuration
CLIP-I ↑
L1 ↓
Cat. ↓
Base: uniform CFG
0.8625
0.1104
152
Base: text gating only
0.9303
0.0653
12
Incumbent: uniform CFG
0.9217
0.0725
–
Incumbent: image anchor only
0.9274
0.0680
–
Incumbent: text gating only
0.9346
0.0615
13
Incumbent: full ASG ( ghi=3.0 )
0.9380
0.0586
8
Table 2: Branch ablations on the 528-turn MagicBrush development split. “Base” is the public IP2P checkpoint; “incumbent” is the finetuned backbone used for the main results. “Cat.” counts turns whose CLIP-I gap to the no-op reference falls below −0.10 .
(a) Anchor dose: MagicBrush (528 turns)
ghi
CLIP-I ↑
L1 ↓
New cat. ↓
1.50
0.9346
0.0615
0
2.25
0.9367
0.0598
0
3.00
0.9380
0.0586
0
4.50
0.9350
0.0592
3
(b) ASG minus reference (311 turns)
Table 3: Dose selection and spatial controls. Panels (a,b) use the MagicBrush development split; new catastrophic turns are counted relative to text-only gating ( ghi=1.5 ), and (b) uses the 311 turns with source–target CLIP-I ≤0.97 . Panel (c) reports PIE-Bench++ controls relative to ASG.
Figure 3: Selected same-seed development comparisons against uniform CFG (7.5,1.5) . Our outputs replace the remote with pizza while retaining facial appearance (a), and recolor the scarf while better retaining its placement (b). Target images provide reference edits.
City University of Hong Kong (Dongguan), Guangdong, China · City University of Hong Kong, Hong Kong, China · Mohamed bin Zayed University of Artificial Intelligence, Masdar, Abu Dhabi