Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we construct a new benchmark, ScribbleEdit, that evaluates the ability of image editing models to perform image editing conditioned on scribbles. This task requires both a deep understanding of the intention of the scribble and an accurate interpretation of its spatial information. In ScribbleEdit, we design an automated data construction pipeline and introduce a dedicated evaluation protocol that explicitly measures intention alignment. Our analysis reveals that existing VLM/LLM-based editing models fail to accurately capture scribble intentions. To guide future progress on scribble-only image editing, we propose a simple yet effective soft-token baseline, which enhances the model's understanding of scribble semantics and outperforms standard image editing models on our benchmark. Our evaluation and baseline together provide a concrete foundation for assessing and improving the scribble-driven image editing.
Figures & tables
Figure 1 : Illustration of image editing driven solely by user-provided scribbles. The edited results are generated by the proposed SSE (Soft-Token Scribble Editing) baseline.
Figure 2 : The pipeline of data generation
Figure 3 : Example of the final generated scribble examples
Figure 4 : Focused similarity: To evaluate model performance, the edited regions of the generated outputs are compared with those of the ground truth. Specifically, the edited areas are cropped, and SSIM and CLIP-sim are computed on the cropped images.
Acc.
Train
0.996
Val.
0.992
Manual
0.969
Table 1 : Classification
Figure 5 : The framework of Soft-token Scribble Editing
Operation
FLUX
Qwen
BAGEL
Step1X
OmniGen2
Nano Banana
GPT
Fine-tuned
SSE
Intention Accuracy
Remove
0.440
0.360
0.700
0.600
0.020
1.000
0.000
0.860
0.980
Move
0.100
0.100
0.040
0.020
0.040
0.040
0.021
0.700
0.920
Flip
0.120
0.020
0.080
0.000
0.000
0.200
0.020
0.840
0.960
Scale up
0.300
0.620
0.200
0.380
0.100
0.260
0.813
0.440
0.960
Scale down
0.180
0.020
0.300
0.100
0.960
0.160
0.560
0.460
0.860
Average
0.228
0.224
0.264
0.220
0.224
0.329
0.281
0.660
0.936
Table 2 : The results of Intention Accuracy, Focused SSIM and CLIP-sim on five scribble-only editing operations. The best results are highlighted in bold .
Figure 6 : Intention understanding results of Nano Banana, the directly fine-tuned OmniGen2, and the SSE framework. For each subfigure, the left side shows the ground-truth editing intention of the scribbles, while the right side presents the intention predicted from the edited images. “None” means the image remains unchanged.
Intention Accueracy
Focused SSIM
Focused CLIP-sim
Operation
SSE
Single
SSE
Single
SSE
Single
Remove
0.980
0.940
0.395
0.439
0.897
0.926
Move
0.920
0.920
0.584
0.604
0.935
0.944
Flip
0.960
1.000
0.507
0.466
0.943
0.939
Scale up
0.960
0.960
0.888
0.896
0.969
0.973
Scale down
0.860
0.980
0.371
0.396
0.918
0.933
Table 3 : Upper-bound results of OmniGen2 for different editing operations. “Single” denotes OmniGen2 fine-tuned on data from a single operation, meaning the model does not need to distinguish between different scribble intentions. The better results between SSE and Single are highlighted in bold .
Figure 7 : The editing results using different prompts
Figure 8 : Two representative groups of failure cases. The first corresponds to incorrect interpretation of the scribble intention, while the second shows cases where the model failed to modify the image.
Scribble-guided image editing allows users to combine simple scribble annotations with text prompts to specify both where and how an image should be edited, enabling flexible interaction with precise spatial control. However, existing models still exhibit unstable performance under this paradigm, especially in multi-task scenarios. To improve performance, we conduct empirical studies using an open-source editing model and reveal an asymmetry in generalization: instruction-level generalization, including across editing tasks and from single-task to multi-task settings, is more challenging than image-domain generalization, such as from synthetic to real-world images or from mosaicked to regular images. This suggests that the primary bottleneck lies in insufficient learning for diverse editing instructions rather than in the image domain gap. Motivated by this insight, we propose three strategies: (a) a Coverage-then-Realism Curriculum, a two-stage pipeline that first builds large-scale synthetic, instruction-rich data for broad task supervision, then curates a small set of real-world data to refine generation realism; (b) Multi-Task Mosaicking, which constructs multi-task training samples by concatenating single-task examples at nearly zero cost while enabling the learned capability to generalize to non-mosaicked images; and (c) an Edit-Focused Loss, which leverages the changed regions between input and output images in synthetic data to focus training on edited regions, improving both learning efficiency and editing accuracy. With these strategies, we substantially improve both single-task and multi-task scribble-guided editing on the VIBE benchmark, achieving state-of-the-art results. We will publicly release our dataset and model.
Mingyi Xu, Jinpeng Lin, Min Zhou +2
Xiamen University · Taobao & Tmall Group of Alibaba
Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often struggle to communicate both exact spatial layouts and specific semantic details simultaneously. While natural language instructions effectively convey high-level semantics like texture and color, they lack spatial specificity. Conversely, freehand scribbles provide rough spatial boundaries but cannot express detailed visual attributes. Consequently, achieving precise control requires combining both modalities. However, existing models struggle to jointly interpret abstract scribbles alongside text due to a lack of specialized training data. In this work, we introduce ScribbleEdit, a large-scale synthetic dataset designed to bridge this gap by combining natural language instructions with freehand scribble inputs for more accurate, controllable edits. We construct this dataset through a synthetic pipeline that automatically generates source-target image pairs via inpainting, which are then paired with human-drawn scribbles and VLM-generated text instructions. Using ScribbleEdit, we evaluate and finetune both diffusion-based and autoregressive unified multimodal image editing models. Our experiments reveal that while off-the-shelf models struggle with abstract scribble inputs, finetuning on our synthetic dataset significantly improves their ability to generate spatially aligned and semantically consistent edits.
Humans naturally communicate through abstract concepts like "mood". However, current image editing benchmarks focus primarily on explicit, literal commands, leaving abstract instructions largely underexplored. In this work, we first formalize the definition and taxonomy of abstract image editing. To measure instruction-following in this challenging domain, we introduce Entity-Rubrics, a framework that breaks down abstract edits into individual, entity-level assessments and achieves strong correlation with human judgment. Alongside this framework, we contribute AbstractEdit, the first benchmark dedicated to abstract image editing across diverse real-world scenes. Evaluating 11 leading models on this dataset reveals a fundamental challenge: standard architectures struggle to balance intent and preservation, commonly defaulting to under-editing or over-editing. Our analysis demonstrates that driving meaningful improvements relies heavily on integrating advanced LLM text encoders and iterative thinking. Looking forward, our entity-based paradigm can generalize beyond assessment to serve as a reward model, enable models to correctly interpret abstract communication, or highlight specific failures in test-time critique loops. Ultimately, we hope this work serves as a stepping stone toward seamless multimodal interaction, closing the gap between rigid machine execution and the natural, open-ended way humans communicate.
Mor Ventura, Roy Hirsch, Yonatan Bitton +2
1Technion – Israel Institute of Technology · 2Google Research · *Work done during an internship at Google Research.