Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
Figures & tables
Figure 1: SolveEdit : infer, execute, preserve.
Figure 2: Application breadth of SolveEdit . One representative case from each of ten domains. Panels show the source, instruction, evaluator questions, and reference answers.
Figure 3: Overview of SolveEdit , SolveScore , and SolveEdit-Plan . Left: cases are grouped by the information needed to determine a valid edit. Center: atomic criteria specify the required change and the visual state to preserve; SolveScore aggregates their satisfaction. Right: SolveEdit-Plan instantiates the transition and compiles an instruction for an unmodified editor.
Figure 4: Statistics of the 2,728-case SolveEdit benchmark. Left: frequent request terms. Center: coverage across 10 domains and 54 subdomains. Right: cases per domain.
Model
Required
Protected
R↑
D↓
SolveScore ↑
SA ↑
RA ↑
VQ ↑
SA ↑
RA ↑
VQ ↑
Image-to-Video Models
HunyuanVideo-1.5 [ 55 ]
16.9
17.9
22.5
42.2
57.4
41.9
17.3
55.5
7.4
Kling V3 [ 36 ]
37.6
37.1
54.7
64.1
59.7
75.5
39.1
34.0
27.5
Open-Source Models
OmniGen2 [ 58 ]
19.8
16.2
25.4
49.9
64.3
53.4
18.1
45.9
9.7
Table 1: Performance on SolveEdit (%, higher is better). Best SolveScore is bold. Evaluation coverage is reported in Appendix C ; video outputs are scored from their final frames. SA, RA, and VQ denote Semantic Accuracy, Relational Accuracy, and Visual Quality.
Figure 5: Reasoning failures behind visually plausible outputs. Top left: SolveScore across IS, SD, and RD for the three leading models. Top right: application-domain-adjusted diagnostic gaps from IS, pooled across models and separated into required and protected conditions. Bottom: representative outputs that remain visually coherent but violate task-specific spatial, semantic, or rule-based requirements.
Figure 6: Effect of explicit transition inference. Left: GPT-Image-2 controls and component ablations. Middle: GPT-Image-2 gains by dependency level. Right: Direct versus SolveEdit-Plan across fixed generators. Every method uses one final generation.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Source category
Cases
AI-generated inputs
1,375
Real photographs
646
Animation and game captures
377
Film and cinematic frames
330
Total
2,728
Appendix
Table 2: Final source composition of the SolveEdit benchmark.
Evidence family
Typical criteria
Evidence produced
OCR and parsers
Text, symbols, digits, labels, state updates
Recognized string/state and normalized comparison
Grounding and segmentation
Entity presence, count, target zones
Instances, boxes/masks, count and containment
Geometry and topology
Position, alignment, routes, occupancy, assembly
Distances, overlaps, adjacency and graph relations
Preservation and similarity
Protected objects, layout, background, quality
Region-level change and similarity signals
VLM contract audit
Scene interpretation and criterion assessment
Pass/partial/fail/abstain verdict and visible evidence
Appendix
Table 3: Evidence families used for atomic criteria.
Generator
Instruction method
R↑
D↓
SolveScore ↑
Δ
GPT-Image-2
Direct
67.6
23.9
57.0
–
SolveEdit-Plan
75.8
8.6
71.6
+14.6
Qwen-Image-Edit-2509
Direct
32.7
47.5
20.8
–
SolveEdit-Plan
40.5
38.5
26.9
+6.1
HunyuanVideo-1.5
Direct
17.3
55.5
7.4
–
SolveEdit-Plan
23.7
39.7
14.0
+6.6
Appendix
Table 4: Matched one-generation evaluation of SolveEdit-Plan and controls (%).