Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40×. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.
Figures & tables
Figure 1: In response to a user question about a figure, a typical LLM will generate a long text response that is difficult to read and register with the figure. In contrast, FigAct generates a visual presentation that combines text narration, grounded visual elements, and a sequence of visual actions to guide the viewer through the explanation. See Figure 4 for a concrete example.
Figure 2: Overview of FigAct. FigAct generates a presentation through three stages. Planning generates the text explanation into a sequence of beats. Grounding groups the SVG into a hierarchical forest and performs multi-round search with Keep , Expand , and Delete operations to identify the visual evidence for each beat. Acting designs presentation actions to the grounded elements and forms the resulting temporal visual presentation.
Figure 3: Synthetic data construction and statistics. (a) SVG coverage across context-window sizes under different preprocessing settings. (b) Construction of grounding trajectories from semantic SVGs. (c) Statistics of the six task types, reporting sample counts and average presentation- and beat-level quantities.
Figure 4: Example of Generated Visual Presentation. We use Figure 2 as an example. For the complete presentation, see Appendix F .
Plan
Ground
Act
End-to-End
Method
Fact. ↑
Cover. ↑
Comp. ↑
Elem. F1 ↑
Tok. (M) ↓
Search-E ↑
Exec. ↑
NVA ↑
Pres-Q ↑
Proprietary MLLMs
Claude Fable 5.1
90.64
90.80
71.00
43.44
1.12
–
77.06
70.88
79.52
+ FigAct
–
–
98.00 +27.00
79.69 +36.25
0.022 -1.098
85.96
81.32 +4.26
73.62 +2.74
84.45 +4.93
Gemini 3.5 Flash
94.60
85.20
50.00
22.58
0.99
–
78.25
71.58
81.23
+ FigAct
–
–
96.00 +46.00
68.85 +46.27
0.015 -0.975
95.35
96.71 +18.46
73.30 +1.72
87.62 +6.39
Table 1: Main results on real-world scientific figure presentation. Across proprietary and open-weight MLLMs, applying FigAct at inference time consistently improves grounding, action generation, and end-to-end presentation quality while substantially reducing SVG token usage. SFT and RL further improve grounding, search efficiency, and action generation for the Qwen3-VL-8B-based FigAct model.
Setting
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
SFT only
52.67
83.04
69.68
82.47
w/o Rground
60.98
84.15
71.91
81.97
w/o Rsearch
60.05
83.28
70.85
82.05
w/o Rrender
61.88
84.89
69.35
82.78
Full
62.27
84.68
72.25
82.16
Table 2: Ablation of RL rewards.
Figure 5: FigAct-8B remains robust as task complexity increases. FigAct-8B consistently outperforms the base model across presentation length, SVG complexity, and search depth.
Comprehension Accuracy
Text Highlight Prop. Ours 66.7% 75.6% 86.7% 83.3%
Method
Likert Rating
μ±σ
1
2
3
4
5
6
7
Comprehension Ease
Text
6
11
17
22
17
11
6
4.00±1.61
Highlight
3
6
11
19
22
18
11
4.66±1.56
Table 3: Human evaluation results.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Three rules for visual primitive grouping. Each example starts from separate SVG primitives and applies one of three high-confidence grouping rules: coincident fill–outline layers, layered raster renderings, or marker–stroke assembly. Primitives satisfying the corresponding rule are connected through Union–Find and consolidated into a single visual group.
Figure 7: SVG grouping statistics for 50 sampled figures. Left: primitive, group, and top-level root counts for each figure. Right: median node counts at each containment depth among figures reaching that depth.
Type
Action
Parameters
Usage
Illustration
View Control
camera
target=<ids|page>
Frame focal evidence with necessary context.
Attention Control
tint
target=<ids> [color=<color>]
Assign a stable color to grounded evidence.
deemphasize
target=<ids|page> [keep=<ids>]
Recede competing content while preserving context.
reemphasize
target=<ids>
Restore previously deemphasized content.
Relational Guidance
arrow
connector=<connector-id>
Highlight an existing relation or direction.
infer_relation
from=<id>, to=<id> relation=<text>
Add a labeled relation between grounded elements.
Appendix
Table 4: Action space of FigAct. Actions are grouped by function types. For each action, we list its parameters, intended usage, and a representative rendered example. Square brackets denote optional parameters.
Figure 8: Structural complexity of 1,666 real-world scientific figures , measured by the numbers of SVG paths, embedded raster images, and text glyphs. Horizontal axes use a log-binned visualization scale.
Figure 9: Interface used in the human evaluation. Participants view a question-conditioned explanation of a scientific figure and then answer a comprehension question.
Model
View
Elem. F1 ↑
GS@0.5 ↑
GS@0.75 ↑
Search-E ↑
Gemini 3.5 Flash
Context only
60.99
64.47
40.53
89.50
Isolated only
67.09
86.87
58.63
95.18
Dual view
68.85
89.63
59.82
95.35
Qwen3-VL-8B
Context only
17.75
21.87
3.90
63.75
Isolated only
26.42
25.10
16.17
69.73
Dual view
28.35
27.83
16.92
68.26
Appendix
Table 5: Ablation of context and isolated rendering for hierarchical grounding. Both models use FigAct only at inference time without FigAct-specific training. Best results are shown in bold and second-best results are underlined.
Setting
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
SFT only
52.67
83.04
69.68
80.56
GRPO
61.41
84.62
70.83
81.53
Stage-Aware GRPO
62.27
84.68
72.25
82.16
Appendix
Table 6: Comparison between standard GRPO and our stage-aware reward optimization.
G
Elem. F1 ↑
Search-E ↑
Exec. ↑
NVA ↑
Pres-Q ↑
4
61.51
83.38
96.64
71.07
79.88
8
62.27
84.68
97.10
72.25
82.16
16
62.46
84.89
97.18
71.87
82.46
Appendix
Table 7: Effect of rollout group size on FigAct performance.
Dataset
Model
Elem. F1 ↑
Search-E ↑
NVA ↑
Pres-Q ↑
Synthetic
Qwen3-VL-8B
28.75
–
67.37
78.91
+ SFT
61.68
85.38
67.84
78.85
+ SFT + RL
63.83
85.92
78.21
79.85
Real-world
Qwen3-VL-8B
9.85
–
67.92
70.18
+ SFT
60.67
83.04
69.68
82.47
+ SFT + RL
62.27
84.68
72.25
82.16
Appendix
Table 8: Evaluation on held-out synthetic and real-world scientific figures. Synthetic and real-world scores are reported separately because the two datasets differ in distribution and difficulty.
Model
Deemph.
Annot.
Tint
Camera
Arrow
Reemph.
Infer Rel.
Abstract.
Claude
20.7
32.5
6.2
4.9
9.0
13.3
6.6
6.9
Gemini
39.4
21.0
7.4
14.1
7.4
9.2
1.0
0.5
GPT
18.9
45.6
3.7
2.8
11.4
14.5
2.1
0.9
InternVL3.5-8B
1.7
1.3
73.3
16.2
2.9
4.6
0.0
0.0
Qwen3-VL-8B
4.3
0.0
94.0
0.8
0.5
0.4
0.0
0.0
Qwen3.5-9B
25.1
9.8
58.5
1.5
1.1
3.2
0.8
0.0
Appendix
Table 9: Distribution of visual actions generated by different models. All values are percentages (%) of the total generated actions for each model and sum to approximately 100% due to rounding.
Figure 10: Representative failure cases of FigAct-8B on real-world scientific figures. (a) Identifier binding error. The model predicts a near-match but nonexistent SVG identifier, causing the intended visual action to fail and leaving the next beat visually unchanged. (b) Imprecise focus. The model identifies the correct semantic region but selects a region that is too broad, leaving distracting neighboring content in the focused view.
Figure 11: Complete FigAct presentation for the example introduced in Figure 2 .
Figure 12: Qualitative comparison of question-conditioned presentations across models. Rows show different models and columns show three presentation beats for the same figure and question. The target narration is: Beat 1: AU-Localization starts with Activation Momentum, which tracks individual activations uij for the i -th AU on the j -th sample ; Beat 2: Discriminative Score compares the positive and negative ratios, ripos and rineg , and applies the MAX operator to obtain scores s1,…,si ; Beat 3: these scores determine the ordering in Ranked AUs, with selected activations highlighted in blue. The scientific figure used as input is reproduced from Feng et al. (2026) .
Figure 13: Qualitative comparison of question-conditioned presentations across models. Rows show different models and columns show four presentation beats for the same figure and question. The target narration is: Beat 1: The framework operates in two distinct search rounds: Round 1 focuses on text cues like lyrics, while Round 2 focuses on visual cues like the singer ; Beat 2: In Round 1, the agent zooms in and crops the lyrics ’We sing till the lights come alive’, then runs a text search to find matching song info ; Beat 3: In Round 2, the agent crops the singer’s silhouette to perform a reverse image search, gathering visual matches for the performer ; Beat 4: Finally, the agent compares the results from both rounds to confirm the correct event and performance time. The scientific figure used as input is reproduced from Tao et al. (2026) .
Figure 14: Qualitative comparison of question-conditioned presentations across models. The target narration is: Beat 1: A query Q is posed at observed frame 3, which is processed by a frame encoder to create a representation in the memory buffer ; Beat 2: The memory buffer sequences representations of frames 1 through 6, with the query Q interleaved right after frame 3 ; Beat 3: As frame 6 is observed, it is sent to both a frame encoder and an Activation Model to determine if a response should be triggered ; Beat 4: If the Activation Model outputs ’True’ at frame 6, an ’ASSISTANT:’ token is appended to the memory buffer ; Beat 5: The compiled buffer is optionally compressed via Round-Decayed Compression before being processed by the Large Language Model to generate the final response. The scientific figure used as input is reproduced from Wang et al. (2026) .
Figure 15: Qualitative comparison of question-conditioned presentations across models. The target narration is: Beat 1: A query Q is posed at observed frame 3, which is processed by a frame encoder to create a representation in the memory buffer ; Beat 2: The memory buffer sequences representations of frames 1 through 6, with the query Q interleaved right after frame 3 ; Beat 3: As frame 6 is observed, it is sent to both a frame encoder and an Activation Model to determine if a response should be triggered ; Beat 4: If the Activation Model outputs ’True’ at frame 6, an ’ASSISTANT:’ token is appended to the memory buffer ; Beat 5: The compiled buffer is optionally compressed via Round-Decayed Compression before being processed by the large language model to generate the final response. The scientific figure used as input is reproduced from Wang et al. (2026) .