Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40×. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.
Figures & tables
Figure 1: In response to a user question about a figure, a typical LLM will generate a long text response that is difficult to read and register with the figure. In contrast, FigAct generates a visual presentation that combines text narration, grounded visual elements, and a sequence of visual actions to guide the viewer through the explanation. See Figure 4 for a concrete example.
Figure 2: Overview of FigAct. FigAct generates a presentation through three stages. Planning generates the text explanation into a sequence of beats. Grounding groups the SVG into a hierarchical forest and performs multi-round search with Keep , Expand , and Delete operations to identify the visual evidence for each beat. Acting designs presentation actions to the grounded elements and forms the resulting temporal visual presentation.
Figure 3: Synthetic data construction and statistics. (a) SVG coverage across context-window sizes under different preprocessing settings. (b) Construction of grounding trajectories from semantic SVGs. (c) Statistics of the six task types, reporting sample counts and average presentation- and beat-level quantities.
Figure 4: Example of Generated Visual Presentation. We use Figure 2 as an example. For the complete presentation, see Appendix F .
Plan
Ground
Act
End-to-End
Method
Fact. ↑
Cover. ↑
Comp. ↑
Elem. F1 ↑
Tok. (M) ↓
Search-E ↑
Exec. ↑
NVA ↑
Pres-Q ↑
Proprietary MLLMs
Claude Fable 5.1
90.64
90.80
71.00
43.44
1.12
–
77.06
70.88
79.52
+ FigAct
–
–
98.00 +27.00
79.69 +36.25
0.022 -1.098
85.96
81.32 +4.26
73.62 +2.74
84.45 +4.93
Gemini 3.5 Flash
94.60
85.20
50.00
22.58
0.99
–
78.25
71.58
81.23
+ FigAct
–
–
96.00 +46.00
68.85 +46.27
0.015 -0.975
95.35
96.71 +18.46
73.30 +1.72
87.62 +6.39
Table 1: Main results on real-world scientific figure presentation. Across proprietary and open-weight MLLMs, applying FigAct at inference time consistently improves grounding, action generation, and end-to-end presentation quality while substantially reducing SVG token usage. SFT and RL further improve grounding, search efficiency, and action generation for the Qwen3-VL-8B-based FigAct model.
Setting
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
SFT only
52.67
83.04
69.68
82.47
w/o Rground
60.98
84.15
71.91
81.97
w/o Rsearch
60.05
83.28
70.85
82.05
w/o Rrender
61.88
84.89
69.35
82.78
Full
62.27
84.68
72.25
82.16
Table 2: Ablation of RL rewards.
Figure 5: FigAct-8B remains robust as task complexity increases. FigAct-8B consistently outperforms the base model across presentation length, SVG complexity, and search depth.
Comprehension Accuracy
Text Highlight Prop. Ours 66.7% 75.6% 86.7% 83.3%
Method
Likert Rating
μ±σ
1
2
3
4
5
6
7
Comprehension Ease
Text
6
11
17
22
17
11
6
4.00±1.61
Highlight
3
6
11
19
22
18
11
4.66±1.56
Table 3: Human evaluation results.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Three rules for visual primitive grouping. Each example starts from separate SVG primitives and applies one of three high-confidence grouping rules: coincident fill–outline layers, layered raster renderings, or marker–stroke assembly. Primitives satisfying the corresponding rule are connected through Union–Find and consolidated into a single visual group.
Figure 7: SVG grouping statistics for 50 sampled figures. Left: primitive, group, and top-level root counts for each figure. Right: median node counts at each containment depth among figures reaching that depth.
Type
Action
Parameters
Usage
Illustration
View Control
camera
target=<ids|page>
Frame focal evidence with necessary context.
Attention Control
tint
target=<ids> [color=<color>]
Assign a stable color to grounded evidence.
deemphasize
target=<ids|page> [keep=<ids>]
Recede competing content while preserving context.
reemphasize
target=<ids>
Restore previously deemphasized content.
Relational Guidance
arrow
connector=<connector-id>
Highlight an existing relation or direction.
infer_relation
from=<id>, to=<id> relation=<text>
Add a labeled relation between grounded elements.
Appendix
Table 4: Action space of FigAct. Actions are grouped by function types. For each action, we list its parameters, intended usage, and a representative rendered example. Square brackets denote optional parameters.
Figure 8: Structural complexity of 1,666 real-world scientific figures , measured by the numbers of SVG paths, embedded raster images, and text glyphs. Horizontal axes use a log-binned visualization scale.
Figure 9: Interface used in the human evaluation. Participants view a question-conditioned explanation of a scientific figure and then answer a comprehension question.
Model
View
Elem. F1 ↑
GS@0.5 ↑
GS@0.75 ↑
Search-E ↑
Gemini 3.5 Flash
Context only
60.99
64.47
40.53
89.50
Isolated only
67.09
86.87
58.63
95.18
Dual view
68.85
89.63
59.82
95.35
Qwen3-VL-8B
Context only
17.75
21.87
3.90
63.75
Isolated only
26.42
25.10
16.17
69.73
Dual view
28.35
27.83
16.92
68.26
Appendix
Table 5: Ablation of context and isolated rendering for hierarchical grounding. Both models use FigAct only at inference time without FigAct-specific training. Best results are shown in bold and second-best results are underlined.
Setting
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
SFT only
52.67
83.04
69.68
80.56
GRPO
61.41
84.62
70.83
81.53
Stage-Aware GRPO
62.27
84.68
72.25
82.16
Appendix
Table 6: Comparison between standard GRPO and our stage-aware reward optimization.
G
Elem. F1 ↑
Search-E ↑
Exec. ↑
NVA ↑
Pres-Q ↑
4
61.51
83.38
96.64
71.07
79.88
8
62.27
84.68
97.10
72.25
82.16
16
62.46
84.89
97.18
71.87
82.46
Appendix
Table 7: Effect of rollout group size on FigAct performance.
Dataset
Model
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
Synthetic
Qwen3-VL-8B
28.75
–
67.37
78.91
+ SFT
61.68
85.38
67.84
78.85
+ SFT + RL
63.83
85.92
78.21
79.85
Real-world
Qwen3-VL-8B
9.85
–
67.92
70.18
+ SFT
60.67
83.04
69.68
82.47
+ SFT + RL
62.27
84.68
72.25
82.16
Appendix
Table 8: Evaluation on held-out synthetic and real-world scientific figures. Synthetic and real-world scores are reported separately because the two datasets differ in distribution and difficulty.
Model
Deemph.
Annot.
Tint
Camera
Arrow
Reemph.
Infer Rel.
Abstract.
Claude
20.7
32.5
6.2
4.9
9.0
13.3
6.6
6.9
Gemini
39.4
21.0
7.4
14.1
7.4
9.2
1.0
0.5
GPT
18.9
45.6
3.7
2.8
11.4
14.5
2.1
0.9
InternVL3.5-8B
1.7
1.3
73.3
16.2
2.9
4.6
0.0
0.0
Qwen3-VL-8B
4.3
0.0
94.0
0.8
0.5
0.4
0.0
0.0
Qwen3.5-9B
25.1
9.8
58.5
1.5
1.1
3.2
0.8
0.0
Appendix
Table 9: Distribution of visual actions generated by different models. All values are percentages (%) of the total generated actions for each model and sum to approximately 100% due to rounding.
Figure 10: Representative failure cases of FigAct-8B on real-world scientific figures. (a) Identifier binding error. The model predicts a near-match but nonexistent SVG identifier, causing the intended visual action to fail and leaving the next beat visually unchanged. (b) Imprecise focus. The model identifies the correct semantic region but selects a region that is too broad, leaving distracting neighboring content in the focused view.
Figure 11: Complete FigAct presentation for the example introduced in Figure 2 .
Figure 12: Qualitative comparison of question-conditioned presentations across models. Rows show different models and columns show three presentation beats for the same figure and question. The target narration is: Beat 1: AU-Localization starts with Activation Momentum, which tracks individual activations uij for the i -th AU on the j -th sample ; Beat 2: Discriminative Score compares the positive and negative ratios, ripos and rineg , and applies the MAX operator to obtain scores s1,…,si ; Beat 3: these scores determine the ordering in Ranked AUs, with selected activations highlighted in blue. The scientific figure used as input is reproduced from Feng et al. (2026) .
Figure 13: Qualitative comparison of question-conditioned presentations across models. Rows show different models and columns show four presentation beats for the same figure and question. The target narration is: Beat 1: The framework operates in two distinct search rounds: Round 1 focuses on text cues like lyrics, while Round 2 focuses on visual cues like the singer ; Beat 2: In Round 1, the agent zooms in and crops the lyrics ’We sing till the lights come alive’, then runs a text search to find matching song info ; Beat 3: In Round 2, the agent crops the singer’s silhouette to perform a reverse image search, gathering visual matches for the performer ; Beat 4: Finally, the agent compares the results from both rounds to confirm the correct event and performance time. The scientific figure used as input is reproduced from Tao et al. (2026) .
Figure 14: Qualitative comparison of question-conditioned presentations across models. The target narration is: Beat 1: A query Q is posed at observed frame 3, which is processed by a frame encoder to create a representation in the memory buffer ; Beat 2: The memory buffer sequences representations of frames 1 through 6, with the query Q interleaved right after frame 3 ; Beat 3: As frame 6 is observed, it is sent to both a frame encoder and an Activation Model to determine if a response should be triggered ; Beat 4: If the Activation Model outputs ’True’ at frame 6, an ’ASSISTANT:’ token is appended to the memory buffer ; Beat 5: The compiled buffer is optionally compressed via Round-Decayed Compression before being processed by the Large Language Model to generate the final response. The scientific figure used as input is reproduced from Wang et al. (2026) .
Figure 15: Qualitative comparison of question-conditioned presentations across models. The target narration is: Beat 1: A query Q is posed at observed frame 3, which is processed by a frame encoder to create a representation in the memory buffer ; Beat 2: The memory buffer sequences representations of frames 1 through 6, with the query Q interleaved right after frame 3 ; Beat 3: As frame 6 is observed, it is sent to both a frame encoder and an Activation Model to determine if a response should be triggered ; Beat 4: If the Activation Model outputs ’True’ at frame 6, an ’ASSISTANT:’ token is appended to the memory buffer ; Beat 5: The compiled buffer is optionally compressed via Round-Decayed Compression before being processed by the large language model to generate the final response. The scientific figure used as input is reproduced from Wang et al. (2026) .
Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation
Scientific figures are the interface through which research claims are inspected and reused, but final published panels rarely expose the data or plotting code that produced them. Recovering this hidden provenance from pixels is therefore underdetermined. We introduce SciFigure2Code, an AI-reconstructed benchmark that instead evaluates presentation recovery: generating editable Python programs that preserve how a scientific panel is arranged and read. Role-specialized Codex agents generate, execute, visually refine, and audit silver-standard presentation programs that capture geometry, visual hierarchy, encodings, annotations, and typography without claiming to recover original measurements or author source code. This reconstruction-and-audit protocol turns final published panels into auditable reference packages; the resulting resource contains 6,740 reviewed panels and SciFigureBench, a balanced 337-panel test set across 31 chart subtypes, five domains, and three complexity levels. Across 14 zero-shot models in image-only and caption-assisted settings, execution, multi-component layouts, axes, legends, and scientific labels remain weak. Claude Opus 4.7 achieves the highest image-only Overall score, Claude Opus 4.6 leads caption-assisted reconstruction, and two-stage plan-then-code prompting improves Overall for all four tested models. SciFigure2Code provides an auditable testbed for agents that construct editable, visually faithful scientific figure presentations.
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In this paper, we build a benchmark, Diagram-MMU, a multi-modal benchmark designed to assess MLLMs' ability for scientific diagram parsing and understanding. Diagram-MMU features 3.7k curated diagrams and 18.3k human-validated questions across six domains. It evaluates MLLMs on three tasks common in vibe writing workspaces: diagram-to-code parsing, diagram-to-code editing, and diagram question answering, alongside agentic settings per task. The evaluation of 12 MLLMs reveals that diagram-to-code tasks are more challenging than diagram question answering: models can reason well over diagrams but struggle to parse and edit them, underscoring the need for methods to enhance MLLMs' capability in diagram-to-code generation. Under agentic settings, most models improve parsing and editing performance but degrade on question answering, while Claude-4.6 Opus consistently improves across all three tasks. Project Page: https://vi-ocean.github.io/projects/diagram-mmu.
Weihao Bo, Shan Zhang, Yanpeng Sun +7
Nanjing University of Science and Technology · Baidu Inc · AIML, Adelaide University +4