Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
Figures & tables
Figure 1: SolveEdit : infer, execute, preserve.
Figure 2: Application breadth of SolveEdit . One representative case from each of ten domains. Panels show the source, instruction, evaluator questions, and reference answers.
Figure 3: Overview of SolveEdit , SolveScore , and SolveEdit-Plan . Left: cases are grouped by the information needed to determine a valid edit. Center: atomic criteria specify the required change and the visual state to preserve; SolveScore aggregates their satisfaction. Right: SolveEdit-Plan instantiates the transition and compiles an instruction for an unmodified editor.
Figure 4: Statistics of the 2,728-case SolveEdit benchmark. Left: frequent request terms. Center: coverage across 10 domains and 54 subdomains. Right: cases per domain.
Model
Required
Protected
R↑
D↓
SolveScore ↑
SA ↑
RA ↑
VQ ↑
SA ↑
RA ↑
VQ ↑
Image-to-Video Models
HunyuanVideo-1.5 [ 55 ]
16.9
17.9
22.5
42.2
57.4
41.9
17.3
55.5
7.4
Kling V3 [ 36 ]
37.6
37.1
54.7
64.1
59.7
75.5
39.1
34.0
27.5
Open-Source Models
OmniGen2 [ 58 ]
19.8
16.2
25.4
49.9
64.3
53.4
18.1
45.9
9.7
Table 1: Performance on SolveEdit (%, higher is better). Best SolveScore is bold. Evaluation coverage is reported in Appendix C ; video outputs are scored from their final frames. SA, RA, and VQ denote Semantic Accuracy, Relational Accuracy, and Visual Quality.
Figure 5: Reasoning failures behind visually plausible outputs. Top left: SolveScore across IS, SD, and RD for the three leading models. Top right: application-domain-adjusted diagnostic gaps from IS, pooled across models and separated into required and protected conditions. Bottom: representative outputs that remain visually coherent but violate task-specific spatial, semantic, or rule-based requirements.
Figure 6: Effect of explicit transition inference. Left: GPT-Image-2 controls and component ablations. Middle: GPT-Image-2 gains by dependency level. Right: Direct versus SolveEdit-Plan across fixed generators. Every method uses one final generation.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Source category
Cases
AI-generated inputs
1,375
Real photographs
646
Animation and game captures
377
Film and cinematic frames
330
Total
2,728
Appendix
Table 2: Final source composition of the SolveEdit benchmark.
Evidence family
Typical criteria
Evidence produced
OCR and parsers
Text, symbols, digits, labels, state updates
Recognized string/state and normalized comparison
Grounding and segmentation
Entity presence, count, target zones
Instances, boxes/masks, count and containment
Geometry and topology
Position, alignment, routes, occupancy, assembly
Distances, overlaps, adjacency and graph relations
Preservation and similarity
Protected objects, layout, background, quality
Region-level change and similarity signals
VLM contract audit
Scene interpretation and criterion assessment
Pass/partial/fail/abstain verdict and visible evidence
Appendix
Table 3: Evidence families used for atomic criteria.
Generator
Instruction method
R↑
D↓
SolveScore ↑
Δ
GPT-Image-2
Direct
67.6
23.9
57.0
–
SolveEdit-Plan
75.8
8.6
71.6
+14.6
Qwen-Image-Edit-2509
Direct
32.7
47.5
20.8
–
SolveEdit-Plan
40.5
38.5
26.9
+6.1
HunyuanVideo-1.5
Direct
17.3
55.5
7.4
–
SolveEdit-Plan
23.7
39.7
14.0
+6.6
Appendix
Table 4: Matched one-generation evaluation of SolveEdit-Plan and controls (%).
Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal-centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step-by-step planning-by-generation paradigm. In this work, we present EAR, an editing-as-reasoning paradigm that reformulates visual planning as a single-step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion-based models in terms of both pixel-wise fidelity and logical validity. We assess leading proprietary and open-source editing models. The results show that they all struggle in the zero-shot setting, finetuning on basic scales enables remarkable generalization to larger in-domain scales and out-of-domain scales and geometries. However, our best model that runs on high-end hardware fails to match the zero-shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.
Zhimu Zhou, Yanpeng Zhao, Qiuyu Liao +2
1Shanghai Jiao Tong University · 3State Key Laboratory of General Artificial Intelligence, BIGAI · 2Renmin University of China
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexplored. Addressing this gap, we introduce Corrective Sequence Planning (CoSPlan) benchmark, where VLMs must plan a sequence of visual actions from an initial scene to a target scene. CoSPlan evaluates models on their ability to imagine and execute a coherent set of visual steps required to reach the goal (Step Completion). To prevent any shortcuts that simply describe the final scene, we introduce an erroneous action in decision-making, which must be detected (Error Detection) and corrected to reach the goal, enabling a deeper understanding of the task. CoSPlan spans across 4 tasks: maze navigation, block re-arrangement, image reconstruction, and object re-organization. Despite using advanced reasoning strategies such as Chain-of-Thought and Scene Graphs, VLMs struggle on CoSPlan, while still showing promising performance in the text domain. Addressing this, we propose Scene Graph Incremental updates (SGI), a novel training-free method to transform images into `textual' scene graphs, enabling step-by-step reasoning through iterative scene graph refinement. SGI yields an average of ~4.4% improvement on CoSPlan w/ generalization on PlanBench and VQA. Link for solving puzzles on the project page.
Shresth Grover, Priyank Pathak, Akash Kumar +1
UCF Institute of Artificial Intelligence, University of Central Florida (UCF)
While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To probe this challenge, we introduce PaintBench, a dynamically scalable benchmark targeting 20 fundamental precise visual editing operations across four categories: geometric transformation, structural manipulation, color change, and symbolic reasoning. Procedural generation with configurable complexity enables an effectively infinite, contamination-resistant evaluation suite, and deterministic pixel-level evaluation eliminates reliance on bias-prone judge models. Across 11 image editing models, we find overall low performance, with the current highest-performing industry leader scoring only 17.1% (mIoU). Task decomposition reveals especially challenging operation types (geometric transformation, most structural manipulation, formula-based color change) and model-specific specializations. Fine-grained benchmark diagnostics further show performance degradations induced by scene variations in object count, background complexity, color scheme, and edit-region size. To test generalization of PaintBench scores to applied task performance, we create a procedural, deterministic evaluation for data visualization editing (TinyGrafixBench) and find strong linear correlation with PaintBench scores (R2=0.91, p<0.001). Altogether, PaintBench provides a rigorous foundation for measuring and driving progress in precise multimodal visual editing.