MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
Authors: Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, +1 more
Organizations: X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Suzhou Laboratory, Suzhou, China · State Key Laboratory for General Artificial Intelligence, BIGAI, Beijing, China · Shanghai Artificial Intelligence Laboratory, Shanghai, China · Jiangsu Key Lab of Language Computing, Suzhou, China · Shanghai Innovation Institution, Shanghai, China
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.
Figures & tables
Figure 1: Survey results from 25 materials graduate students: tool usage frequency across sub-fields (left) and distribution of primary interaction interface types (right).
Figure 2: A complete XRD diffraction pattern requires cross-tool collaboration: raw XRD data is analyzed in JADE, then plotted and annotated in OriginPro.
Figure 3: Overview of the MatToolBench framework. Tasks are dispatched by a central runner to one of three specialized agents: the GUI Agent operates materials science desktop applications (e.g., JADE, VESTA, Avantage) via screenshot perception and coordinate-based interaction; the Code Agent generates Python scripts to query materials database APIs (e.g., MP, OQMD, OPTIMADE); and the Origin Agent combines GUI interaction with OriginPro script generation for experimental data plotting. All agents execute inside a Windows 11 VM hosted within a Docker container, communicating through a Flask-based control server. Task-specific evaluators assess the final environment state after each episode.
Cat.
Domain
#
E/M/H
Avg
Tot.
GUI
Avantage
20
14/2/4
3.25
65
DM
20
13/5/2
3.35
67
JADE
20
15/5/0
2.95
59
MS
20
0 6/11/3
4.10
82
VESTA
20
13/2/5
3.45
69
Origin
Origin
16
0 0/7/9
2.00
32
Table 1: Benchmark task statistics. E/M/H: easy / medium / hard task counts. Mixed tasks are reported separately as cross-tool diagnostics; total sub-criteria are 454 single-tool criteria plus 30 mixed diagnostic criteria.
Doubao seed-1-8
Kimi k2.5
Claude sonnet-4.6
GPT 5.4
Qwen3-VL 235B
Qwen3-VL 32B
Qwen3-VL 8B
Cat.
Domain
Sc.
SR
Sc.
SR
Sc.
SR
Sc.
SR
Sc.
SR
Sc.
SR
Sc.
SR
GUI
Avantage
44.6
35.0
32.3
20.0
41.3
35.0
32.3
15.0
30.8
10.0
26.2
15.0
18.5
15.0
JADE
54.2
25.0
50.8
25.0
57.6
25.0
50.8
20.0
45.8
10.0
37.3
10.0
37.3
5.0
DM
56.7
20.0
52.2
20.0
40.3
20.0
50.7
20.0
52.2
10.0
6.0
0.0
38.8
10.0
MS
47.6
15.0
35.4
10.0
52.4
20.0
56.1
30.0
43.9
0.0
26.8
0.0
20.7
0.0
VESTA
52.2
20.0
59.4
25.0
58.0
25.0
65.2
20.0
52.2
10.0
29.0
10.0
30.4
10.0
Table 2: Accuracy results on MatToolBench . Sc. (Score): normalized task score (%) averaged over sub-criteria; SR: success rate (%, all sub-criteria satisfied). Bold : best average metric within each category; bold italic : best overall average metric.
Type
Parameter
Value
Condition
Script
Hint
Env Setup
Injected Content
hint
full
—
✓
—
GUI workflow instructions
GUI
gui_hint_mode
no_hint
baseline
—
×
—
generic guidelines only
hint
full
—
✓
—
API examples & docs
Code
code_hint_mode
no_hint
baseline
—
×
—
generic guidelines only
script + hint
full
✓
✓
✓
template + workflow
no_script + hint
partial
×
✓
✓
workflow only
Table 3: Ablation conditions for each task type. Env Setup indicates whether the task environment automatically opens Code Builder (Alt+4) before the episode starts. In no_hint mode this step is suppressed, requiring the agent to discover the workflow tool independently.
Figure 4: Ablation study: SR comparison between hint (solid) and no_hint (dashed) prompt modes for Doubao-seed-1-8 (blue) and GPT-5.4 (orange). Left: GUI tasks. Right: Code tasks. Shaded bands indicate the gap between hint and no_hint conditions.
Figure 5: Origin three-level ablation for Claude Sonnet 4.6. Bars show deterministic success points (dark purple, max 16 pts) and Aesthetic Score (light purple, max 16 pts) for each condition. Total scores and percentages are annotated at bar ends.
Score
SR
Domain
Acc.
Prec.
Rec.
F1
Acc.
Prec.
Rec.
F1
JADE
1.00
1.00
0.99
0.99
1.00
1.00
1.00
1.00
MS
1.00
0.99
1.00
0.99
1.00
1.00
1.00
1.00
Avantage
0.99
0.98
0.98
0.98
1.00
0.98
1.00
0.99
VESTA
0.98
0.94
0.99
0.97
0.99
0.94
1.00
0.97
DM
0.99
0.99
0.94
0.96
1.00
0.98
1.00
0.99
Table 4: GUI evaluator reliability under an n×n cross-evaluation protocol ( n=20 ). DM uses the filtered audit described in Appendix I.
Figure 6: Radar chart comparing Origin figure quality across four task types (Raman, XRD, XPS, FTIR). Subplots are task types; axes are visual correctness, aesthetic quality, and task completeness (1–5). Dashed lines are means from 9 humans; solid lines are four LLM judges.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Task distribution across 10 domains and their sub-categories.
Module
Method
Parameters
Description
mouse
move_id(id)
Element ID (int)
Optional; available only when SoM or accessibility-tree observations are enabled. Not used in the reported screenshot-only experiments
move_abs(x, y)
Normalized coords ∈[0,1]
Move cursor to a relative screen position; (0,0) =top-left, (1,1) =bottom-right
single_click()
—
Left single-click at the current cursor position
double_click()
—
Left double-click at the current cursor position
right_click()
—
Right-click at the current cursor position
scroll(dir)
"up" / "down"
Scroll the active window by 400 units in the given direction
Appendix
Table 6: Complete Computer action interface and execution constraints on the Windows VM. Calls are sent as Python strings via HTTP POST and executed server-side with exec(code, {"computer": computer}) .
Table 9: Bundled sample data files pre-installed in the evaluation VM, grouped by domain.
Domain
Getter Type
Description
Avantage
check_multiple_files_imported
Confirms that multiple XPS scan channels ( e.g. , C 1s, O 1s, Zn 2p) are loaded.
check_dialog_exists
Detects the presence of a specific dialog window via OCR.
check_dialog_parameter
Extracts and validates a parameter value inside a dialog.
check_stacked_graph
Verifies that spectra are displayed in stacked mode.
check_background_added
Confirms background curve addition to a spectrum.
check_energy_axis_reversed
Validates that the binding-energy axis is inverted.
Appendix
Table 10: Getter types used in GUI-based domains, all paired with exact_match .
Domain
exact_match
detect_kv_match
detect_file_match
Avantage
✓
–
–
DM
✓
–
–
JADE
✓
–
–
MS
✓
–
–
VESTA
✓
–
–
Origin
✓
–
–
Appendix
Table 11: Metric functions used per domain. ✓: all tasks use this metric; fractions indicate partial usage.
JADE
MS
Avantage
VESTA
DM
Task
F1
SR-F1
F1
SR-F1
F1
SR-F1
F1
SR-F1
F1
SR-F1
in1
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
0.86
1.00
in2
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
0.80
0.86
in3
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
in4
1.00
1.00
1.00
1.00
0.88
0.92
1.00
1.00
1.00
1.00
in5
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
Appendix
Table 12: Per-task evaluator reliability across five GUI domains: Score F1 and SR-F1. Values below 1.00 are explicitly reported. Means are macro-averaged over the 20 tasks in each domain: JADE 0.991, MS 0.994, Avantage 0.980, VESTA 0.979, DM 0.936.
Doubao seed-1-8
Kimi k2.5
Claude sonnet-4.6
GPT -5.4
Qwen3-VL 235B
Qwen3-VL 32B
Qwen3-VL 8B
Cat.
Domain
Steps
Steps*
FTR
Steps
Steps*
FTR
Steps
Steps*
FTR
Steps
Steps*
FTR
Steps
Steps*
FTR
Steps
Steps*
FTR
Steps
Steps*
FTR
GUI
Avantage
36.1
19.9
25.0
42.4
36.0
60.0
41.1
26.1
25.0
39.2
7.8
50.0
39.7
8.0
71.4
46.2
31.7
60.0
44.9
42.4
66.7
JADE
34.6
24.2
50.0
35.3
15.6
44.4
35.9
15.2
44.4
29.0
8.7
85.7
41.8
19.5
71.4
42.9
5.1
85.7
48.7
41.4
0.0
DM
38.3
27.6
42.9
42.9
30.8
40.0
45.4
28.6
0.0
39.3
27.3
57.1
38.1
18.5
77.8
50.0
NaN
NaN
50.0
50.0
NaN
MS
47.6
42.0
50.0
48.8
39.0
50.0
45.1
42.1
50.0
41.6
34.9
37.5
50.0
NaN
100.0
50.0
NaN
NaN
50.0
NaN
NaN
VESTA
39.1
22.3
50.0
41.1
36.6
57.1
41.1
36.6
57.1
30.3
40.8
83.3
33.5
10.5
75.0
44.1
13.3
33.3
50.0
50.0
NaN
Appendix
Table 13: MatToolBench efficiency metrics. Steps: mean steps across all episodes (max 50; fewer is better). Steps*: mean steps on successful (SR = 1) episodes. FTR (%): false termination rate. Cat. : task category. ‘NaN’: no successful episodes (Steps* undefined).
Judge
Avg C/A/T
Parsed
Pearson
95% CI
Spearman
GPT-5.4
4.33/4.15/4.38
43/43
0.513
[-0.180, 0.826]
0.348
Doubao-seed-1-8
4.61/4.43/4.80
43/43
0.677
[-0.017, 0.896]
0.402
Gemini-3.1-pro
4.60/4.25/4.68
43/43
0.694
[0.033, 0.891]
0.346
Qwen3-VL-235B
4.81/4.46/4.92
43/43
0.737
[0.104, 0.928]
0.518
Appendix
Table 14: Origin aesthetic-judge validation. C/A/T denote visual correctness, aesthetic quality, and task completeness. All judges parse all 43 images.
Figure 9: Representative Origin figures generated by Claude Sonnet 4.6 with script+hint (top) and no_script+no_hint (bottom). Scores are computed as s=(C+A+T)/(3×5) . Template scripts produce high-quality outputs ( s≥0.98 ), whereas GUI-only execution reduces quality ( s≤0.70 ) or fails ( s=0 ).
Figure 10: Execution trajectory and scoring sub-criteria for the Avantage task 8ebfc15b (hard, 7 pts). Each panel shows an intermediate GUI state validated by a domain-specific getter, including OCR, parameter extraction, file-existence, and application-state checks. The evaluator scores all sub-criteria independently at episode end, and the final task score is the fraction of passed criteria.
Figure 11: Representative visual-state detection methods used by the GUI evaluators. Panels (a)–(b) identify the active VESTA rendering style from the vertical position of blue radio-button pixels. Panel (c) detects the yellow-highlighted polyhedral icon using HSV thresholding. Panel (d) uses structural similarity against a reference image to validate the target crystal orientation. Panels (e)–(f) count green control handles in DigitalMicrograph to confirm ellipse and rectangle creation. Each detection result is paired with exact_match to produce a binary sub-criterion score.
Tool
Domain
Description
JADE
characterization
MDI JADE is a commercial XRD analysis suite for phase identification, Rietveld refinement, and diffractogram comparison using the PDF card database.
Avantage
characterization
Thermo Fisher Avantage is the standard software for XPS spectrum processing, including peak fitting, quantification, and elemental binding-energy analysis.
VESTA
Visualisation
VESTA (Visualisation for Electronic and STructural Analysis) renders crystal structures in 3-D; tasks cover bond/polyhedra style changes, CIF import, and orientation control.
DigitalMicrograph
characterization
Gatan DigitalMicrograph (DM) is the de-facto TEM image analysis platform; tasks cover FFT, line profiles, annotation, and drawing-tool operations.
Materials Studio
Simulation
Dassault BIOVIA Materials Studio is a molecular modelling environment; tasks require building supercells, setting force-field parameters, and running geometry optimisations.
OriginPro
Analysis
OriginLab OriginPro is a scientific graphing and statistics package; tasks span curve fitting, multi-panel figure layout, axis formatting, and batch export via Code Builder (Python).
Appendix
Table 15: Overview of the ten software tools and APIs covered by MatToolBench .
Figure 12: GUI grounding error in an Avantage comparison-view task (Doubao-seed-1-8). Each misclick on a visually similar toolbar icon opens a new independent data grid instead of selecting the target layout option.
Figure 13: Premature termination in an Avantage stacked-graph task (Doubao-seed-1-8). The agent mistakes the pre-highlighted “Display Modes” tab for evidence of task completion after only 6 of ≤ 50 steps.
Figure 14: Deadlock in a Materials Studio file-import task (Doubao-seed-1-8). The agent alternates clicks between two adjacent left-panel entries every step, never entering the target subdirectory; the step budget expires with the dialog still open.
Figure 15: Domain knowledge gap in a JADE profile-fitting task (Doubao-seed-1-8). The agent correctly completes the fitting but fails at the final export step because it lacks knowledge that the VM has no virtual PDF printer installed.