Reliable infographic generation requires facts, symbols, and visual relations to remain consistent through rendering and revision. Correcting one element also requires tracking its supporting evidence and the dependencies affected by the change. We present \textbf{InfoAgent}, a training-free framework for \emph{evidence-bound visual-symbolic program synthesis}. Its Infographic Visual Description (IVD) records factual payloads, evidence provenance, execution routes, and verification obligations in a typed dependency graph. Retrieved design priors guide compilation, and layered execution combines raster synthesis with editable symbolic and binding objects while retaining their traces. Dependency-aware repair localizes corrections, rechecks affected dependencies, and requires protected obligations to remain satisfied under the declared checkers. Unresolved obligations remain explicit. On IGenBench, InfoAgent achieves 93.0 Q-ACC and 59.0 I-ACC. We also introduce InfoGraphicBench-Evidence, where complete-checklist pass rates on 200 test requests increase from 21.5% for Same-IVD Prompt to 23.5% for the initial layered output and 28.5% after repair, using the same evidence and initial IVD. On 120 audited repair cases, localized repair edits 12.4% of the canvas on average, compared with 67.3% for global regeneration.
Figures & tables
Figure 1: Overview of InfoAgent. A user request, retrieved evidence, and design priors are compiled into a typed IVD. Each information-bearing element retains evidence provenance, a spatial region, an execution route, binding dependencies, and checkable obligations. Visual, symbolic, and binding layers are composed with execution traces. IVD-guided verification then maps element-indexed violations to scoped content, symbolic, binding, layout, or visual patches.
Method
Q-ACC ↑
I-ACC ↑
Out ↑
Benchmark-reported references Tang et al. (2026)
NanoBanana-Pro Google (2025)
90.0
49.0
–
Seedream-4.5 ByteDance (2025)
61.0
6.0
–
GPT-Image-1.5 OpenAI Team (2025)
55.0
12.0
–
Controlled baselines
Direct T2I Wu et al. (2025)
43.0
2.0
100.0
Table 1: Reliability on the complete IGenBench benchmark.
Method
Out ↑
ReqCov ↑
Full ↑
NonSup. (%) ↓
ReqText-F1 ↑
Bind-Q ↑
Q-Align ↑
Direct T2I Wu et al. (2025)
100.0
76.74
17.5
1.63
82.7
4.53
4.508
RAG Prompt
100.0
79.40
19.5
1.18
83.3
4.56
4.515
Same-IVD Prompt
100.0
81.00
21.5
1.02
86.4
4.58
4.522
Gen-Searcher Feng et al. (2026b)
100.0
82.10
22.5
0.87
85.4
4.61
4.556
GenClaw Ye et al. (2026)
98.0
81.50
25.0
0.76
92.7
4.60
4.048
LLM-to-SVG/HTML
96.5
79.61
23.0
0.80
94.7
4.58
3.920
Table 2: Controlled results on 200 InfoGraphicBench-Evidence test requests. Evidence-conditioned methods use the same evidence bundle; Direct T2I is evidence-free.
Variant
CritPass ↑
Fact ↑
Layout-Q ↑
ReqText-F1 ↑
Bind-Q ↑
Q-Align ↑
Full InfoAgent
61.5
91.6
4.45
96.5
4.64
4.621
InfoAgent w/o repair
46.0
88.0
4.30
92.4
4.47
4.502
Flat plan
42.5
86.8
4.13
88.1
4.24
4.389
Symbolic content to raster
38.5
90.1
4.43
85.2
4.55
4.529
w/o evidence binding
49.0
84.2
4.44
92.9
4.58
4.515
w/o binding route
45.0
89.9
4.42
93.0
4.19
4.511
Table 3: Ablations on InfoGraphicBench-Evidence. ReqText-F1, Bind-Q, and Q-Align are shared with Table 2 ; CritPass, Fact, and Layout-Q isolate critical reliability, required-fact preservation, and layout quality.
Strategy
Fix@ Det
Repair Rec.
New Img.
Collat. Reg.
Area Edit
Localized repair
85.9
77.8
3.3
0.8
12.4
w/o dependency closure
79.1
71.7
7.5
4.9
7.8
Global regeneration
74.8
67.8
10.0
6.8
67.3
Table 4: Repair effectiveness and locality on 120 audited repair-eligible outputs.
Figure 2: Reliability breakdown and scoped-repair dynamics. (a) Category-wise Q-ACC on IGenBench. (b) Mean detected critical violations and cumulative edited area over three repair rounds on 120 audited repair-eligible outputs.
Figure 3: Qualitative comparison on financial-event and product-specification requests. Boxes mark factual, symbolic, binding, and hierarchy differences; ✓ , × , and ? denote satisfied, violated, and partially satisfied requirements. All outputs are shown without manual correction.
Table S5: Three-generation robustness. Values are mean ± standard deviation.
Planner
Executor
ReqCov
Full
ReqText-F1
Bind-Q
Q-Align
Gemini 3.1 Pro
Same-IVD
81.00
21.5
86.4
4.58
4.522
Gemini 3.1 Pro
InfoAgent
83.54
28.5
96.5
4.64
4.621
Claude Opus 4.6
Same-IVD
80.4
20.0
85.7
4.55
4.49
Claude Opus 4.6
InfoAgent
82.8
27.0
95.8
4.61
4.60
Appendix
Table S6: Planner transfer on the complete test set.
Evaluator
Full κ
Fact F1
Binding F1
Gemini 3.1 Pro
0.76
91.4
87.2
Claude Opus 4.6
0.73
90.8
86.5
Human–human
0.79
93.0
89.1
Appendix
Table S7: Agreement with blinded human evaluation.
Checker
N
Prec.
Rec.
F1
Unknown
Rendering-tree
240
97.7
95.0
96.3
0.0
Evidence support
150
92.5
84.1
88.1
5.3
Binding geometry
110
90.8
79.8
84.9
2.7
Visual semantic
100
89.9
76.5
82.7
10.0
Appendix
Table S8: Checker audit on 600 obligation decisions.
Threshold
Coverage
Risk
Critical Unknown
False repair
Relaxed
97.0
14.5
3.0
5.8
Default
94.2
9.8
5.8
3.3
Conservative
85.0
6.2
15.0
1.7
Appendix
Table S9: Coverage–risk trade-off for model-assisted verification.
Figure S1: Additional IGenBench results. Each row presents the benchmark prompt, reference infographic, and InfoAgent output. The examples cover treemap, bubble-chart, pictogram, and line-chart layouts and are shown without manual correction.
Figure S2: Additional evidence-controlled comparisons. The examples compare factual coverage, symbolic fidelity, local grounding, and visual hierarchy across generation methods, and include representative unresolved and unsuccessful-repair cases.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Hong Kong Baptist University +1