Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.
Figures & tables
Figure 1: Editing comparison. Pixel-space warping provides explicit but imperfect motion evidence. MoRe-Drag injects this evidence into the generative trajectory, enabling the editor to follow the specified motion while restoring natural structure and appearance.
Figure 2: Conceptual comparison with existing drag-based and text-based editing paradigms. Existing drag methods use warped image as an intermediate result, leading to accurate but unnatural edits, while text-based editors produce natural results without precise drag control. Our method uses warped image as motion evidence to achieve both accurate dragging and natural appearance.
Figure 3: Overview of MoRe-Drag . (a) Region-aware recomposition: the evidence and original image are decomposed into role-specific regions, producing role masks that guide latent recomposition, so motion cues are injected into the corresponding regions. (b) Generation-stage adaptive conditioning: at early timesteps (t≥tc) , the model uses the warped image with a fixed prompt to enforce motion consistency; at later timesteps (t<tc) , it switches to the original image with the user/auto prompt for semantic and appearance refinement.
Figure 4: Toy editing from different initial timesteps. Top: starting latent timestep; bottom: LPIPS to the original image.
Figure 5: Instruction interaction and data construction. Top: drag annotation construction from ByteMorph . Bottom: LoRA-based MLLM fine-tuning with AR loss, enabling automatic instruction inference from drag annotations.
Model
Params
DragBench-SR
DragBench-DR
MD ↓
CP ↑
PF ↑
MD ↓
CP ↑
PF ↑
Test-time Optimization-based Methods
DragDiffusion
2.1B
41.08±0.43
92.75±0.38
77.15±1.49
31.61±0.23
96.46±0.43
78.45±0.95
DragNoise
2.1B
47.64±1.08
79.00±0.25
61.56±0.31
30.05±0.01
91.46±1.34
81.02±0.18
GoodDrag
2.1B
28.94±0.40
94.50±0.13
76.50±0.39
23.21±0.12
95.36±0.24
83.94±0.04
Optimization-free Methods
Table 1: Results on DragBench-SR/DR . CP and PF are scaled by 100. Metrics of Qwen-Image-Edit and LongCat-Image-Edit are not highlighted, as they are not drag-based methods. “+ FLUX.1-Fill-dev” denotes an Inpaint4Drag variant using a stronger inpainting model. (*) indicates results from the original papers, as the models were not publicly available at the time of evaluation. In the Params column, (-) indicates that MoRe-Drag keeps the generative backbone frozen.
Figure 6: Qualitative comparisons with state-of-the-art drag-based editing models. Blue and red points denote handle and target points, respectively. LongCat-Image-Edit is a text-based editing method; Inpaint4Drag and MoRe-Drag take the warped image as input, while the other methods are driven by point inputs. We provide zoomed-in views of the edited regions for the two warped-image-based methods to highlight local structural fidelity. More results are shown in Fig. 11
Figure 8
Figure 9: Instruction interaction examples. Compared with vanilla instruction inference, LoRA adaptation produces drag-aware instructions that better match the visual annotations, reducing unintended artifacts and hallucinated content.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Main models
Key settings
Size
Raw dataset
-
-
780K
Motion-edit filtering
Qwen3.5-9B
local human/object motion; no global camera change
Table 4: Summary of the drag-instruction dataset construction pipeline.
Figure 10: Illustration of bidirectional warping and mask decomposition. Given handle–target drag annotations, bidirectional warping produces a warped observation that provides pixel-space motion evidence. We decompose the result into preservation, refinement, and inpainting regions for region-aware latent recomposition: blue denotes the anchor preservation Mp , yellow denotes the refinement mask Mr , and red denotes the inpainting mask Mi .
Figure 11: Additional qualitative comparisons on diverse drag-based editing examples. MoRe-Drag faithfully follows the specified drag motion while preserving object structure, texture details, and background consistency.
Setting (RAR steps, GAC switch)
DragBench-SR
MD ↓
CP ↑
PF ↑
(6, 2)
37.00±0.48
92.75±0.00
69.72±0.28
(11, 9)
21.09±0.20
95.75±0.25
80.63±0.38
(16, 12)
18.54±0.22
96.36±0.38
86.36±4.06
(21, 14)
18.43±0.31
95.87±0.88
83.50±2.36
(30, 20)
17.98±0.37
94.25±1.00
80.75±1.61
Appendix
Table 5: Hyperparameter analysis on DragBench-SR .
MD ↓
CP ↑
PF ↑
Vanilla
19.94±0.27
95.12±0.13
78.63±1.38
w/ LoRA
18.54±0.22
96.36±0.38
86.36±4.06
Appendix
Table 6: Quantitative analysis of instruction interaction on DragBench-SR .
Figure 12: Additional comparisons of drag-aware instruction inference. Each column shows a drag-annotated input, the instruction inferred by the vanilla MLLM, and the instruction inferred by the LoRA-adapted MLLM. LoRA adaptation improves grounding to the visual drag annotations and produces instructions that better reflect the intended object motion or deformation.
Figure 13: Failure cases. When the warped observation contains severe errors or ambiguous disocclusions, the injected motion evidence may not faithfully represent the intended edit. In these cases, MoRe-Drag may favor plausible restoration over the desired motion, such as incomplete head manipulation or unsuccessful lid articulation.
Figure 14: Failure cases of drag-aware instruction inference. The LoRA-adapted MLLM may still infer imperfect instructions when drag annotations are too close to each other, especially for fine-grained shape changes.
Step k
4
8
16
20
24
28
tc
0.87
0.73
0.47
0.33
0.20
0.07
MD ↓
21.01±0.1
19.27±0.25
18.59±0.32
17.98±0.09
17.52±0.24
17.36±0.16
CP ↑
96.09±0.63
97.10±0.13
97.47±0.25
97.10±0.13
95.96±0.50
96.46±0.26
PF ↑
83.33±2.68
85.10±1.99
80.08±1.29
81.06±0.51
79.31±0.74
80.94±0.13
Appendix
Table 7: Effect of switching timestep tc on DragBench-SR .
Figure 15: Comparison of base editor reconstruction ability. Under null conditions, LongCat-Image-Edit reconstructs the source images more faithfully than Qwen-Image-Edit, with better preservation of identity, structure, and appearance. This stronger reconstruction stability benefits motion-evidence injection and non-edited region preservation in MoRe-Drag .
Method
Sampling steps
Avg. time / image (s)
Peak GPU memory (GB)
Baseline editor
30
30.74
16.05
MoRe-Drag
30
41.31
17.18
Appendix
Table 8: Runtime and memory comparison. We report the average generation time per image and peak GPU memory over 10 samples.
City University of Hong Kong (Dongguan), Guangdong, China · City University of Hong Kong, Hong Kong, China · Mohamed bin Zayed University of Artificial Intelligence, Masdar, Abu Dhabi