cs.AISep 29, 2026

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Authors: Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, +1 more

Organizations: X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Suzhou Laboratory, Suzhou, China · State Key Laboratory for General Artificial Intelligence, BIGAI, Beijing, China · Shanghai Artificial Intelligence Laboratory, Shanghai, China · Jiangsu Key Lab of Language Computing, Suzhou, China · Shanghai Innovation Institution, Shanghai, China

Abstract

Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

    May 28, 2026Wanhao Liu, Jiaqing Xie, Qian Tan +10Materials ScienceMaterials

  2. CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

    Sep 30, 2026Chenmu Zhang, Levi Felix, Jun-Jie Zhang +6Materials ScienceArtificial Intelligence Agents

  3. TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

    May 16, 2026Zhiqiang Liu, Wenhui Dong, Yilang Tan +3Multimodal AgentsAgentic Benchmarks