Authors: Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, +4 more
Organizations: University of Sydney · Microsoft Research · Fudan University · Nankai University · Shanghai Jiao Tong University · Tsinghua University · University of Science and Technology of China · University of Waterloo
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
Figures & tables
Figure 1: VibeEdit enables image editing without requiring a separate text prompt. Users express editing intent directly on the image through canvas instructions , including spatial marks such as circles, scribbles, and drag gestures, optionally accompanied by short text notes . Given the resulting canvas-annotated image , VibeEdit supports diverse editing operations, including object addition, removal, replacement, attribute modification, and movement. More visualizations are provided in Figure 6 in the Appendix.
Figure 2: Supervised learning edit-pair construction. A shared pipeline performs object selection, edit prompt synthesis, and segmentation. Operation-specific pipelines then generate source–target pairs for addition, removal, attribute modification, replacement, and movement. Canvas annotations are shown for illustration and are rendered online during training.
Figure 3: VibeEdit training pipeline. Left: the vision–language encoder processes annotated image A , while the VAE separately encodes source image I and rasterized instruction Cˉ . The target image supervises region-weighted training. Right: the base model is refined with DiffusionNFT using rewards for edit success, outside preservation, and local edit quality.
Method
Add V/P
Remove V/P
Attribute V/P
Replace V/P
Move V/P
Overall V/P
Visual-only baselines
Qwen-Image-Edit-2511
17.7 / 17.0
22.1 / 14.3
28.1 / 17.4
21.1 / 16.5
10.0 / 16.1
19.8 / 16.3
LongCat-Image-Edit
8.8 / 15.0
39.0 / 18.9
29.5 / 14.8
27.8 / 16.2
10.5 / 13.5
23.1 / 15.6
FLUX.2-klein-9B
30.0 / 24.1
34.3 / 27.6
56.9 / 26.8
43.3 / 25.7
23.8 / 27.9
37.6 / 26.4
FireRed-Image-Edit-1.0
26.9 / 17.9
35.4 / 17.1
28.6 / 13.4
19.0 / 16.6
4.2 / 14.6
22.8 / 16.0
JoyAI-Image-Edit
69.1 / 23.9
38.1 / 27.8
49.9 / 24.1
63.0 / 26.0
12.0 / 20.9
46.4 / 24.5
Table 1: Main comparison on the VibeEdit benchmark. We tailor the input format to different model types. Visual-only baselines receive the canvas-annotated image A ; text-instructed baselines receive the source image I and a text instruction. VibeEdit uses (I,C) without a separate text prompt. Entries report V (VLM rubric score (%)) / P (outside-region PSNR (dB)). Bold and underlined values indicate the best and second-best results.
Table 5
Figure 4: Qualitative comparison across all five edit types without separate text prompts. Baseline methods exhibit typical errors such as instruction-mark retention, target-instance confusion, unintended changes to unrelated regions, and inaccurate object relocation. Additional qualitative results are provided in Figure 6 in the Appendix.
Table 7
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Training Data
Add
Remove
Attribute
Replace
Move
Overall
Add only
42.7
45.0
62.7
46.6
28.8
45.2
Remove only
28.7
81.4
40.7
50.7
37.6
47.8
Attribute only
60.1
35.0
59.8
55.4
19.1
45.9
Replace only
61.2
53.4
60.1
77.6
27.9
56.1
Move only
23.8
31.3
44.7
34.3
55.9
38.0
All five tasks
67.6
73.3
66.0
72.9
59.1
67.8
Appendix
Table 6: Effect of task composition in supervised learning. All models are evaluated on the full benchmark. Entries report VLM rubric scores (%); higher is better. Bold and underlined values indicate the best and second-best results in each column.
Table 9
Canvas Instruction
VLM ↑
PSNR ↑
X mark
58.4
27.0
X mark w/ text
78.9
30.8
Scribble
55.4
27.6
Scribble w/ text
81.4
36.1
Appendix
Table 9: Comparison of removal instruction designs. Four models are independently trained on the same removal data under supervised learning and evaluated on the removal subset. All designs circle the target; “w/ text” adds “remove it” to the X mark or scribble.
Figure 5: Four removal instruction designs and outputs from their corresponding base models. Each design uses a circled X mark or scribble, with or without a short handwritten text cue.
Pooled rubric pass rates
Edit Model
Human (%)
GPT-5.6-sol (%)
VibeEdit
76.6
77.0
Agreement with majority-vote human labels
Judge
Accuracy / F1 (%) ↑
Cohen’s κ↑
GPT-5.6-sol
93.0 / 95.5
0.805
Appendix
Table 10: Validation of GPT-5.6-sol against majority-vote human labels on 100 randomly sampled VibeEdit outputs. Pass rates are pooled over 947 matched binary questions. Accuracy, F1, and Cohen’s κ measure question-level agreement.
Attention
Modulation
MLP
Trainable Params.
VLM ↑
PSNR ↑
✓
377.5M
64.1
26.6
✓
✓
707.8M
67.8
27.0
✓
✓
✓
943.7M
60.3
22.7
Appendix
Table 11: Effect of the DiT modules adapted with LoRA. All configurations use the same rank, learning rate, training data, and optimization budget.
Operation
Spatial Marks
Rasterized Note
Addition
Circle around the insertion region
add {object}
Removal
Circle and scribble over the object
remove it
Attribute Modification
Circle around the selected object
change {attribute} to {value}
Replacement
Circle around the selected object
replace to {description}
Movement
Source and destination circles connected by a directed arrow
None
Appendix
Table 12: Canvas instruction grammar for each operation. Braced placeholders are filled from the structured edit fields. Text templates reproduce the renderer’s wording.
Resource
Edit Category
Count
Avg. Words
Supervised Learning Corpus
Addition/Removal
435,496
–
Attribute Modification
458,154
–
Object Replacement
393,440
–
Object Movement
266,972
–
Total
1,554,062
–
RL Set
Each Edit Type
704
–
Appendix
Table 13: Composition of the supervised learning corpus, RL set, and VibeEdit benchmark. Average word counts refer to the external text prompts used for text-instructed benchmark baselines.
Figure 6: Additional VibeEdit results across the five edit types. Each example pairs a canvas-annotated image with the generated edit.
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic intent. To unify the strength of textual and visual prompts, we present Text-Vision Co-Instructed Image Editing, which jointly models textual instructions as semantic intent and sparse visual instructions as spatial guidance, aiming to achieve precise and intent-faithful image manipulation. To this end, we first construct a textual-visual instruction paired dataset with more than 23K samples derived from dynamic videos, enabling aligned supervision for cross-modal instruction. We then propose TV-Edit, a Textual-Visual instruction unified Editing framework to contextualize drag or point-based visual instructions with image-text semantics and lift them into semantic-aware control representations for pretrained editing backbones. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. Finally, we establish TV-Edit-Bench, a deliberately designed benchmark to evaluate semantic faithfulness, spatial alignment, and visual consistency with ground-truth references and controlled textual-visual variations for reliable assessment. Our experiments across multiple editing backbones demonstrate that TV-Edit consistently yields more precise and intent-faithful edits, significantly outperforming state-of-the-art instruction-based and drag-based baselines.
Chenxi Xie, Yuhui Wu, Qiaosi Yi +1
The Hong Kong Polytechnic University · OPPO Research Institute
Instruction-guided image editing has a training-time blind spot. Generative editors are never required to semantically verify whether their outputs actually satisfy the instruction. Supervision stops at reconstruction and input textual-level conditioning. This produces incomplete edits, spatial spillover, and poor localization. We present IABEdit, a model-agnostic framework that embeds differentiable semantic verification into training. A frozen vision-language model extracts spatially-aware descriptors from the ground-truth edit. A trainable aligner then reproduces them from the generated output. The residual between the two becomes a gradient that teaches the generator both what to edit and where, with no inference-time VLM cost. IABEdit is compatible with diverse backbones, including U-Net (Stable Diffusion) and MMDiT (FLUX), without altering their inference pipelines. On MagicBrush, it improves structural fidelity by +3.49 DINO-I over the best diffusion baseline and +1.26 over the best overall baseline, while remaining competitive on instruction alignment. It also achieves state-of-the-art instruction adherence performance on RealEdit and EMU Edit benchmarks based on embedding-based metrics. Most consequentially, on the D-LORD surveillance benchmark, it surpasses the proprietary Gemini agent by +5.13 DINO-P under heavy occlusion, where preserving identity is hardest. This shows that gradient-aligned VLM distillation holds up under real-world-like surveillance and occlusion conditions. Human and GPT-4o evaluations confirm perceptually precise, well-localized edits.
Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. As the first to explore training-free image editing with Generative Refinement Networks (GRN), we observe that its refinement process is inherently suitable for editing and offers a promising way to address these limitations. Motivated by this observation, we introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on the GRN. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be revised as the image evolves. More specifically, RefineEdit combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking. Bit routing starts from an intermediate source state and uses signed probability differences between the two branches to identify editable positions and bits. It directs selected bits toward editing refinement while anchoring the rest to the evolving source trajectory. Adaptive spatial freezing limits unnecessary expansion of the editing region, while finite bit locking maintains recent bit activations to support continued editing. The overall framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE, and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods. Code is available at https://github.com/mura1n/RefineEdit.
Yulong Chen, Ziqian Zhang, Haoyu Zhang +4
City University of Hong Kong (Dongguan), Guangdong, China · City University of Hong Kong, Hong Kong, China · Mohamed bin Zayed University of Artificial Intelligence, Masdar, Abu Dhabi