cs.GROct 7, 2026

ScribbleEdit: A Benchmark for Scribble-Only Image Editing

Authors: Jie Ren, Hao Kang, Kai Guo, Yiding Yang, Bo Liu, Liming Jiang, Qing Yan, Zichuan Liu, +4 more

Organizations: MIT · ByteDance · Michigan State University

Abstract

Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we construct a new benchmark, ScribbleEdit, that evaluates the ability of image editing models to perform image editing conditioned on scribbles. This task requires both a deep understanding of the intention of the scribble and an accurate interpretation of its spatial information. In ScribbleEdit, we design an automated data construction pipeline and introduce a dedicated evaluation protocol that explicitly measures intention alignment. Our analysis reveals that existing VLM/LLM-based editing models fail to accurately capture scribble intentions. To guide future progress on scribble-only image editing, we propose a simple yet effective soft-token baseline, which enhances the model's understanding of scribble semantics and outperforms standard image editing models on our benchmark. Our evaluation and baseline together provide a concrete foundation for assessing and improving the scribble-driven image editing.

Figures & tables

Explore similar work

CardsList
  1. Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking

    May 25, 2026Mingyi Xu, Jinpeng Lin, Min Zhou +2Image EditingInstruction

  2. ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text

    May 1, 2026Anya Ji, George Ma, Téa Wright +4Diffusion-Based Image EditingSynthetic Data

  3. Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

    May 14, 2026Mor Ventura, Roy Hirsch, Yonatan Bitton +2Image EditingMulti-Modal Data