InfoEdit: Probing Global Layout Reasoning in Infographic Editing
Authors: Cheng Yang, Chufan Shi, Huijuan Wang, Bo Shui, Yaokang Wu, Muzi Tao, Yibo Yan, Xuezhe Ma, +1 more
Organizations: University of California San Diego · University of Southern California · University of Illinois Urbana-Champaign · Carnegie Mellon University
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to this global layout reasoning capability as reflow. Existing image-editing benchmarks neither provide a dedicated setting for structured visual content nor evaluate the reflow capability. We introduce InfoEdit, a novel benchmark of 1,000 infographics across eight logical-relation families, paired with 4,000 editing instructions across four editing tasks, and a reflow-aware evaluation protocol. Across eight frontier editors, only GPT-Image-2 clears 60% average success rate; most models fall below 7%, and no editor exceeds 36% on the Swap-Block task even with perfect target localization. We further show that code-level editing can match the strongest pixel-level editor, revealing complementary strengths across tasks. InfoEdit identifies reflow as a central challenge in structured visual content editing and provides a diagnostic benchmark to facilitate future progress.
Figures & tables
Figure 1: A typical “vibe-design” workflow with structured visual content. A user asks an image editor to add a missing step to an infographic. The instruction names one element, but the surrounding nodes and connectors must all reflow for the result to be usable. InfoEdit measures whether current editors can perform this reflow reliably.
Figure 2: Overview of InfoEdit . We curate 1,000 infographics covering eight logical-relation families (left), annotate four editing tasks for each infographic (middle), and evaluate model outputs on two binary aspects, edit compliance and content preservation (right). Data examples are provided in Appx. A.3 .
Figure 3: Distributional summary of InfoEdit . (a) Distinct logical-relation families per infographic; 19% contain one, 81% contain two to six. (b) Fraction of infographics containing each family; an infographic can contain multiple families. (c) Source aspect-ratio distribution over eight ratios.
Benchmark
Task
Domain
#Samples
Logical Relation
Reflow- Aware
Evaluation Metric
Image Editing
ImgEdit Ye et al. (2025)
Image Editing
Natural
811
✗
✗
MLLM-judge
WiseEdit Pan et al. (2025)
Image Editing
Natural
1,220
✗
✗
MLLM-judge
GEditBench-v2 Jiang et al. (2026)
Image Editing
Natural
1,200
✗
✗
Pairwise MLLM
Infographics Generation
BizGen Peng et al. (2025)
Text-to-Image
Infographic
1,000
✗
–
MLLM + OCR
Table 1: Comparison of InfoEdit with related benchmarks. “–” indicates not applicable to text-to-image generation. The 4,000 samples come from 1,000 infographics with 4 edit tasks each.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
Proprietary
GPT-Image-2
99.0
86.7
85.9
85.4
80.7
72.0
29.0
79.2
25.5
96.3
65.1
64.4
62.0
Nanobanana-Pro
87.1
64.0
57.8
58.8
57.8
40.2
36.5
79.8
32.0
87.2
36.9
36.4
41.6
Nanobanana-2
85.1
58.6
52.6
65.3
49.4
39.0
30.0
72.1
24.0
92.0
36.4
36.3
38.0
Nanobanana
23.2
23.8
9.2
3.1
14.6
0.4
1.6
41.8
0.9
25.4
8.4
3.5
3.5
Table 2: Editing performance of eight image-editing models on InfoEdit across four editing tasks. EC = edit compliance, CP = content preservation, SR = success rate, the conjunction of both. All values are percentages. Avg. is the mean SR across tasks.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
Code
Gemini-3.1-Pro
85.7
72.3
67.4
85.6
70.7
68.0
68.8
77.6
59.6
69.4
45.2
45.1
60.0
Gemini-3.5-Flash
83.3
68.1
64.1
84.1
64.5
62.0
60.1
79.2
54.8
70.4
42.4
42.2
55.8
Gemini-3.1-Flash
79.6
65.5
60.6
63.4
49.1
40.2
56.1
72.7
47.1
38.0
5.3
5.2
38.3
Code + Image
Table 3: Code-level editing on InfoEdit . Code feeds the source code alone; code + image adds the rendered image. EC = edit compliance, CP = content preservation, SR = the conjunction of both. All values are percentages. Avg. is the mean SR across tasks.
Model
Hint
EC
CP
SR
Δ SR
GPT-Image-2
w/o
29.0
79.2
25.5
–
w/
34.2
92.7
32.9
+7.4
Nanobanana-Pro
w/o
36.5
79.8
32.0
–
w/
39.4
83.6
35.2
+3.2
Nanobanana-2
w/o
30.0
72.1
24.0
–
w/
34.2
79.1
27.5
+3.5
Table 4: Effect of interactive selection hints on Swap-Block performance.
Figure 4: Success rates across difficulty levels for edit tasks: Expand-Text , Insert-Element , and Swap-Block .
Figure 5: Distribution of failure types per task. We categorize unsuccessful cases by the diagnostic sub-field they violate, separated into Edit Compliance (blue) and Content Preservation (purple). Detailed error cases in Appx. K .
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
Family
#Templates
Family
#Templates
List
16
Matrix
6
Process
16
Pyramid
6
Relationship
15
Hierarchy
4
Cycle
7
Picture
9
Total
79
Appendix
Table 5: Templates per logical-relation family in the logical template library. Each template is implemented in both HTML and PowerPoint.
Figure 6: List : Tab List and Trapezoid List.
Figure 7: Process : Timeline and Step Down Process.
Figure 8: Cycle : Basic Cycle and Basic Radial.
Figure 9: Hierarchy : Top-Down Hierarchy and Horizontal Hierarchy.
Figure 10: Relationship : Basic Venn and Nested Target.
Figure 11: Matrix : Basic Matrix and Titled Matrix.
Figure 12: Pyramid : Pyramid List and Inverted Pyramid.
Figure 13: Picture : Circle Picture Grid and Vertical Picture List.
Figure 14: Distribution of infographic topics across 31 categories in InfoEdit , demonstrating broad coverage of real-world domains.
Task
Magnitude Bucket
#Samples
Ratio
Difficulty
Expand-Text
5–10 words
285
28.50%
Easy
11–15 words
302
30.20%
Medium
16–20 words
205
20.50%
Medium
21–25 words
99
9.90%
Hard
26–30 words
109
10.90%
Hard
Insert-Element
1 element
170
17.00%
Easy
Appendix
Table 6: Difficulty-level statistics for editing tasks with explicit magnitude axes. Counts are reported over 1,000 instructions per task.
Figure 15: Example of the Expand-Text task in InfoEdit . The editor must lengthen the body text of a specified card while preserving its content. The expected edit reflows nearby elements to maintain readability and layout consistency, whereas the unexpected edit inserts the longer text without properly adjusting the local layout.
Figure 16: Example of the Insert-Element task in InfoEdit . The editor must insert new nodes into an existing logical structure. The expected edit places them at the specified positions and adjusts the surrounding structure, whereas the unexpected edit adds the nodes without sufficient reflow, resulting in an invalid local layout.
Figure 17: Example of the Swap-Block task in InfoEdit . The instruction asks the editor to exchange multiple logical blocks on the canvas. The expected edit swaps the specified blocks as whole units while preserving their internal content and adapting the surrounding layout; the unexpected edit only partially satisfies the requested swaps or damages the original layout structure.
Figure 18: Example of the Reshape-Canvas task in InfoEdit . The instruction asks the editor to re-render a portrait infographic as a landscape infographic. The expected edit globally reflows the content to fit the target aspect ratio while preserving the original information and visual style; the unexpected edit changes the canvas shape but fails to perform a complete global re-layout.
Figure 19: Interactive selection hint for Swap-Block . Matching numeric labels indicate the two blocks to be exchanged. The boxes are used only for target localization and should not appear in the edited output.
Output model
Samples
Agreement
GPT-Image-2
800
0.920
Nanobanana-2
800
0.934
HunyuanImage-3.0
800
0.981
Appendix
Table 7: Agreement between Gemini-3.1-Pro and the human-majority verdict across output models.
Axis
Subset
Agreement
Task
Expand-Text
0.920
Insert-Element
0.905
Swap-Block
0.925
Reshape-Canvas
0.930
Difficulty
Easy
0.943
Medium
0.912
Appendix
Table 8: Gemini-3.1-Pro agreement with the human majority on GPT-Image-2 outputs, stratified by task, difficulty, and predicted outcome.
Judge
GPT-Image-2
Nanobanana-2
Gemini-3.1-Pro
0.920
0.934
Claude-Opus-4.7
0.914
0.928
Appendix
Table 9: Agreement with the human-majority verdict for two MLLM judges.
Output model
Agreement over five runs
GPT-Image-2
0.924±0.007
Nanobanana-2
0.936±0.008
Appendix
Table 10: Repeated Gemini-3.1-Pro evaluation against human-majority labels.
Family
GPT-Image-2
Nanobanana-Pro
Nanobanana-2
List
78.9
45.3
40.5
Matrix
75.2
37.0
34.5
Process
77.1
47.3
45.0
Picture
75.5
36.5
35.2
Cycle
49.3
20.6
31.6
Pyramid
72.7
49.5
46.5
Appendix
Table 11: Insert-Element SR by host logical-relation family. All values are percentages.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
Source
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
SR
GPT-Image-2
Overall
99.0
86.7
85.9
85.4
80.7
72.0
29.0
79.2
25.5
96.3
65.1
64.4
62.0
HTML
99.1
86.3
85.5
85.8
80.3
71.9
29.5
78.9
25.9
96.6
64.6
64.0
61.8
PPT
98.5
88.5
87.5
84.0
82.5
72.5
27.0
80.5
24.0
95.0
67.0
66.0
62.5
Nanobanana-2
Overall
85.1
58.6
52.6
65.3
49.4
39.0
30.0
72.1
24.0
92.0
36.4
36.3
38.0
HTML
84.9
58.1
52.1
65.9
48.9
38.9
30.5
71.8
24.4
92.3
35.9
35.9
37.8
Appendix
Table 12: Pixel-level editing results by source format.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
Source
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
SR
Gemini-3.1-Pro
Overall
88.2
72.0
68.5
85.3
72.2
71.0
69.0
78.4
60.2
71.8
47.2
46.8
61.6
HTML
96.5
79.6
77.6
92.6
80.9
79.5
75.6
85.4
68.0
75.9
53.4
52.9
69.5
PPT
55.0
41.5
32.0
56.0
37.5
37.0
42.5
50.5
29.0
55.5
22.5
22.5
30.1
Gemini-3.5-Flash
Overall
85.1
69.5
65.8
85.4
67.2
63.2
60.9
79.8
56.8
71.2
46.6
43.1
57.2
HTML
93.1
76.9
74.5
92.8
75.6
70.8
66.8
86.9
64.1
75.1
52.8
48.8
64.5
Appendix
Table 13: Code-level editing with code and image input, broken down by source format.
Model
P → L
L → P
S → L
S → P
GPT-Image-2
69.1
59.4
72.5
67.5
Nanobanana-Pro
36.7
31.8
65.0
60.0
Nanobanana-2
41.2
32.0
35.0
37.5
Appendix
Table 14: Reshape-Canvas SR by transformation direction. P, L, and S denote portrait, landscape, and square.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
Setting
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
SR
GPT-Image-2
Original
99.0
86.7
85.9
85.4
80.7
72.0
29.0
79.2
25.5
96.3
65.1
64.4
62.0
Color-Change
98.2
86.2
84.8
84.2
81.3
74.2
31.0
81.2
27.2
98.2
67.1
66.8
63.3
Nanobanana-2
Original
85.1
58.6
52.6
65.3
49.4
39.0
30.0
72.1
24.0
92.0
36.4
36.3
38.0
Color-Change
86.2
57.2
51.7
63.2
49.0
37.2
32.1
73.2
24.8
92.3
35.2
34.9
37.2
Appendix
Table 15: Color-palette control. All values are percentages.
Expand-Text
Insert-Element
Swap-Block
Reshape-Canvas
Avg.
Model
Setting
EC
CP
SR
EC
CP
SR
EC
CP
SR
EC
CP
SR
SR
GPT-Image-2
Original
99.0
86.7
85.9
85.4
80.7
72.0
29.0
79.2
25.5
96.3
65.1
64.4
62.0
Chinese
93.0
84.5
82.1
80.4
74.2
67.2
25.2
75.2
24.1
93.2
61.2
57.2
57.7
Nanobanana-2
Original
85.1
58.6
52.6
65.3
49.4
39.0
30.0
72.1
24.0
92.0
36.4
36.3
38.0
Chinese
86.2
62.1
56.4
42.5
46.2
25.3
22.6
67.2
15.2
86.2
28.2
26.3
30.8
Appendix
Table 16: Preliminary Chinese translation pilot. All values are percentages.
Model Name
Type
Inference Interface
Proprietary models
GPT-Image-2
API
gpt-image-2
GPT-Image-1.5
API
gpt-image-1.5
Nanobanana-Pro
API
gemini-3-pro-image-preview
Nanobanana-2
API
gemini-3.1-flash-image-preview
Nanobanana
API
gemini-2.5-flash-image
Appendix
Table 17: Models used in our experiments and their inference interfaces. For proprietary models, we report the API model identifier used in our calls; for open-weight models, we report the corresponding Hugging Face repository.
Figure 20: Misplacement after Insert-Element. The three new circles were inserted into the Core Demographics radial diagram, but “Marine” is placed in the wrong position. It is incorrectly positioned on the opposite side of the diagram between “Military” and “Space”.
Figure 21: Style change after Insert-Element. The five new segments were inserted between “Cloud Ops” and “Mobile UI”, but they do not match the style of the existing segments. Instead of being integrated as arc segments within the main ring, they are rendered as smaller circular nodes branching outward below the cycle.
Figure 22: Incomplete content after Insert-Element. The Core Economic Drivers row is incomplete. There are only three of the four requested cards were inserted, while “Advanced Manufacturing” is missing between “Cleantech Energy” and “Silicon Engineering”.
Figure 23: Element loss after Expand-Text. The “Bulk pricing” was successfully expanded with the longer description, but the adjacent “Site delivery” card in the top-right of the Chapter 04 “Core strategy” grid has been removed. It makes the grid incomplete and disrupts the original 3×2 layout.
Figure 24: Style change after Expand-Text. The “4.5 Stories” caption was correctly expanded, but the “Broadcasting Spire” label at the top of “Structural Hierarchy” pyramid has changed from black text on a light background to white text. This breaks visual consistency with the original styling.
Figure 25: Element loss after Swap-Block. The swap was only partially executed. The “Transfer Cycle” was moved to the middle-left position previously occupied by “Core Academic Pillars” but the “Core Academic Pillars” is missing entirely instead of being relocated to the bottom-middle.
Figure 26: Style change after Swap-Block. The “Strategic” and “Global Integration” modules were swapped in position, but their original styling was not preserved. The ascending size progression of the roadmap steps is lost.
Figure 27: Structural break after Reshape-Canvas. The “Supply ecology” network diagram has been structurally altered. The connection topology between the nodes no longer matches the original.
Figure 28: Structural break after Reshape-Canvas. All original content was preserved when rendered as a portrait infographic, but the block ordering does not follow a coherent reading flow.
Figure 29: Style change after Reshape-Canvas. The infographic was re-rendered into a landscape canvas with all original content preserved, but the card styling has been altered.
Figure 30: Text rewritten after Swap-Block using Qwen-Image-Edit. The model partially changes the layout, but many original texts, labels, and chart annotations are rewritten or corrupted instead of being preserved. This shows that the model fails to maintain non-target content.
Figure 31: Prompt for MLLM-as-a-Judge evaluation on Expand-Text editing tasks.
Figure 32: Prompt for MLLM-as-a-Judge evaluation on Insert-Element editing tasks.
Figure 33: Prompt for MLLM-as-a-Judge evaluation on Swap-Block editing tasks.
Figure 34: Prompt for MLLM-as-a-Judge evaluation on Reshape-Canvas editing tasks.