EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
Authors: Tianfu Wang, Leilei Ding, Ziyang Tao, Yi Zhan, Zhiyuan Ma, Wei Wu, Yuxuan Lei, Junyang Wang, +7 more
Organizations: Hong Kong University of Science and Technology (Guangzhou) · University of Science and Technology of China · Peking University · Hong Kong University of Science and Technology
High-fidelity diagram creation requires the complex orchestration of semantic topology, visual styling, and spatial layout, posing a significant challenge for automated systems. Existing methods also suffer from a representation gap: pixel-based models often lack precise control, while code-based synthesis limits intuitive flexibility. To bridge this gap, we introduce EvoDiagram, an agentic framework that generates object-level editable diagrams via an intermediate canvas schema. EvoDiagram employs a coordinated multi-agent system to decouple semantic intent from rendering logic, resolving conflicts across heterogeneous design layers. Additionally, we propose a design knowledge evolution mechanism that distills execution traces into a hierarchical memory of domain guidelines, enabling agents to retrieve context-aware expertise adaptively. We further release CanvasBench, a benchmark consisting of both data and metrics for canvas-based diagramming. Extensive experiments demonstrate that EvoDiagram exhibits excellent performance and balance against baselines in generating editable, structurally consistent, and aesthetically coherent diagrams. Our code is available at https://github.com/AuraX-AI/EvoDiagram.
Figures & tables
Figure 1 : Comparison of diagram generation paradigms. Unlike pixel-based generation (limited control) or code-based synthesis (high barrier), Canvas-based creation unifies AI actionability with human-interpretable UI editing, bridging the representation gap.
Figure 2 : The overview of EvoDesign framework. (a) Agentic Creation System : A multi-agent pipeline where specialized agents for structure, style, and layout coordinate via a shared symbolic schema, followed by a closed-loop refinement agent to resolve cross-layer conflicts. (b) Hybrid Experience Search : A structural exploration of the design space using both vertical refinement and horizontal comparison. (c) Design Knowledge Distillation : A hierarchical process that distills execution traces into specialized domain guidelines and universal design principles for storage in the design expertise Memory.
Figure 3 : The dataset construction pipeline and the overview of CanvasBench.
Category
Method
Content
Visual
Cognitive
CCF ( ↑ )
CCL ( ↑ )
CSR ( ↑ )
VVA ( ↑ )
VSC ( ↑ )
VFC ( ↑ )
VSB ( ↑ )
GCE ( ↑ )
GCA ( ↑ )
GSE ( ↑ )
Diffusion
GPT-4o-Image
2.200
3.110
2.710
3.990
4.354
4.153
3.923
4.301
2.751
3.043
NanoBanano
1.986
2.871
2.657
3.767
3.838
3.533
3.671
3.761
2.665
2.990
Flux.2 flex
1.871
2.782
2.643
3.836
3.868
3.614
3.704
3.780
2.712
2.922
L a T e X Tikz
GPT-5.2
4.096
3.933
4.067
2.833
3.257
2.933
2.543
3.295
3.171
2.505
Gemini-3-Pro
4.201
4.129
4.086
3.368
3.670
3.641
3.124
3.871
3.416
2.541
Table 1 : Main results on CanvasBench. Performance is measured across content, visual, and cognitive dimensions.
Figure 4 : Visual comparison of different baseline methods on a representative diagramming task.
Table 3: Styling-related properties and categorical parameter spaces of tldraw shape elements
Category
Property
Restricted Parameter Space
Position
x, y
Positive Integers
Dimensions
w, h
Positive Integers
Rotation
rotation
Integer Degrees
Alignment
align
start , middle , end
Layering
index
Sequential Rank Index
Appendix
Table 4: Layout-related properties and discretized parameter spaces of tldraw shape elements
Table 5 : Detailed Taxonomy for Diagram Image Retrieval. The taxonomy intersects 21 structural types with 30 semantic domains.
Figure 5 : Information density by diagram type. The box plots illustrate the distribution of character counts for each category, revealing high semantic variation and significant textual depth across the dataset.
Figure 6 : Hierarchical distribution of vertical domains. The sunburst chart depicts the balanced coverage across six primary disciplines and 30 granular sub-domains, ensuring the benchmark tests generalizability across diverse knowledge fields.
Font Consistency, Line & Shape Consistency, Color Theme Consistency
Flow Coherence
VFC
Reading Path & Navigation
Main Reading Direction, Low Path Ambiguity, Key Path Trackability
Appendix
Table 6 : Overview of Evaluation Metrics
Figure 8 : A case study of refinement iterations.
Table 7 : Example of sample strategy ( Ks ) within the design memory M .
Table 8 : Examples of domain guideline ( Ks ) within the design memory M .
Table 9 : Examples of general principle ( Kp ) within the design memory M .
Figure 9 : Human intuitively refines the generated diagram via UI-friendly operations in our web application. The interface facilitates a fluid transition from agentic generation to manual manipulation.
Scientific diagrams are essential for communicating complex methodologies in academic papers. A natural way for researchers to specify such diagrams is through rough sketches, where text labels, connectors, and spatial arrangements express early semantic and topological intentions. However, sketches are usually incomplete, making them insufficient for directly producing publication-quality diagrams. Existing sketch-based generation methods mainly reconstruct the sketch itself, while recent text-driven diagram generation frameworks rely on textual semantics and do not fully exploit the topological structure contained in sketches. In this paper, we introduce DiagramRAG, a lightweight retrieval-augmented framework for sketch-based scientific diagram completion. Given a user sketch, DiagramRAG retrieves reference diagrams that are both semantically relevant to the sketch content and topologically compatible with its structure, and uses them to guide downstream diagram generation. To enable efficient structure-aware retrieval, we represent diagrams as knowledge graphs, synthesize sketch variants at different simplification levels, and train an embedding model to align sketches with compatible diagrams in a shared space. The retrieved references further provide content, topology, and visual priors for completing and rendering the final diagram. Experiments show that DiagramRAG achieves F1-scores of 0.848 and 0.802 on DiagramBank and FigureBench, respectively, and improves generation quality with the best VLM-as-a-Judge score of 7.170, while reducing inference latency to 35.48 seconds per sample. Our code and data are available at https://anonymous.4open.science/r/DiagramRAG-A262 and https://huggingface.co/datasets/anonymous-review-a262/DiagramSketch.
Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer, rather than only the region containing the final response. Creating such structured evidence across diagrams, charts, maps, circuits, and infographics is time-consuming, and existing annotation tools tightly couple their interfaces to dataset-specific formats. We present DIAGRAMS, a lightweight, schema-driven review framework that decouples interface logic from dataset-specific JSON structures through an internal meta-schema and dataset adapters. Given an image and QA pair with optional candidate regions, the system performs QA-conditioned evidence selection and proposes the regions required for reasoning. When QA pairs or candidate regions are missing, it generates them and supports human verification and refinement. Across six Diagram QA datasets, model-suggested evidence achieves 85.39% precision and 75.30% recall against reviewer-final selections (micro-averaged). These results indicate that the review-first framework reduces manual region creation while maintaining high agreement with final reasoning-level attributions. We release a public demo and installable package to support dataset auditing, grounded supervision creation, and grounded evaluation.
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri +4
Arizona State University · University of Maryland · IIITDM +1
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.
Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7
University of California, Los Angeles · Google · University of Illinois Urbana-Champaign +1