Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.
Figures & tables
Figure 1: Overview of CompMat-Bench, illustrated with research steps from a study of isotope effects in boron arsenide ( Wu, 2026 ) . For an expensive step, input preparation and output analysis are separate tasks; prepared inputs are graded without running the simulation, and analysis starts from our precomputed outputs. Lighter calculations are executed by the agent. Tasks are evaluated singly or as a workflow of connected steps, with full or reduced method guidance. Fixed rules grade submissions against hidden references from our reproduction.
Published
Expensive
Excluded from
Benchmark
Domain
studies
computation
evaluation
Grading
FIRE-Bench ( Wang et al., 2026b )
ML
✓
✗
–
LLM judge
ReplicationBench ( Ye et al., 2025 )
astrophysics
✓
(✓)
(✓)
fixed rules
PaperBench ( Starace et al., 2025 )
ML
✓
(✓)
✗
LLM judge
AutoMat ( Huang et al., 2026a )
materials
✓
✓
(✓)
LLM judge
SimulCost ( Cao et al., 2026 )
physics sim.
✗
(✓)
✗
fixed rules
Table 1: CompMat-Bench and prior agent benchmarks on three properties: tasks derived from published studies, tasks involving expensive computation, and expensive computation excluded from evaluation; the last column gives the grading method. ✓ yes, ✗ no, (✓) in part, – no expensive computation in the tasks.
Figure 2: A single task and a workflow from a study of magnons in antiferromagnetic CrN ( Kim et al., 2026 ) . Both panels show the same three stages. Left: in single task 013 (SF, SR), the first two stages (dashed) are precomputed, and the agent computes the spin waves from their outputs. Right: in workflow C06 (WF, WR), the agent runs all three stages. Shaded prompt text is removed in the reduced-guidance version, and the rest of the prompt is identical. Bottom: the hidden reference value and tolerance for one graded quantity.
Figure 3: Composition of the 94 tasks. Left: tasks by computational method. Right: tasks by the scientific code they involve. Shading gives the kind of task, as in Fig. 1 .
Pass rate (%)
Prepare
Analyze
Run and
All
API cost
Model
Harness
inputs
outputs
analyze
tasks
(USD)
n=30
n=29
n=35
n=94
DeepSeek V4.1 Flash
MatClaw
36.7
79.3
80.0
66.0 (62/94)
3.86
Gemini 3.8 Flash
MatClaw
70.0
96.6
100.0
89.4 (84/94)
106.43
GPT-5.6 Sol
MatClaw
76.7
96.6
97.1
90.4 (85/94)
27.00
Table 2: Pass rates of four agents on the 94 single tasks with full guidance, one run per task. API cost is the cost of the 94 runs at each vendor’s standard list price; the Codex CLI run was billed by subscription and has no recorded cost.
Tasks by passing runs
Model
Pass rate (%)
5/5
4/5
3/5
2/5
1/5
0/5
DeepSeek V4.1 Flash
62.0
7
3
2
3
3
2
Gemini 3.8 Flash
88.0
16
1
1
0
1
1
GPT-5.6 Sol
93.0
17
2
0
0
0
1
Table 3: Repeatability of three agents on a selected subset of 20 single tasks with full guidance, five runs per task. The pass rate is over the 100 runs, and the remaining columns count the tasks by how many of the five runs passed.
Figure 4: Pass rates of three agents under the four evaluation conditions of Section 3.2 . Each cell gives the pass rate and the count behind it. The single-task column covers the 52 tasks whose reduced prompt differs from the full prompt, and the workflow column covers the eight workflows with three runs each. The Δ cells give differences in percentage points: reduced minus full guidance (bottom row) and workflow minus single task (right column, different task sets).
Figure 5: Failure analysis: failed runs of three agents by cause under the four evaluation conditions of Section 3.2 . DeepSeek, Gemini and GPT denote DeepSeek V4.1 Flash, Gemini 3.8 Flash and GPT-5.6 Sol. Each cell gives the number of failed runs and their share of all runs in that row.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
No.
Study
Journal
Month
Methods
Graph steps
Selected
Tasks
1
( Li et al., 2025 )
Phys. Rev. B
2025-09
DFT, Structure
10
4
7
2
( Jiang et al., 2026 )
Nano Lett.
2026-01
Model Hamiltonians, DFT
7
5
5
3
( Xie et al., 2026 )
Phys. Rev. B
2026-01
DFT, Structure, Spin models
10
5
8
4
( Devaraj et al., 2026 )
Phys. Rev. B
2026-02
DFT, Model Hamiltonians, Structure
12
5
7
5
( Hofer et al., 2026 )
Nano Lett.
2026-02
Model Hamiltonians
6
3
4
6
( Parlak et al., 2026 )
Phys. Rev. B
2026-02
Phonons, DFT
10
3
6
Appendix
Table A1: The 15 source studies. Month is the publication month. Methods lists the method families of Fig. 3 for the tasks of each study, shortened to their first words. Graph steps counts the research steps in the full graph of each study, Selected counts the steps in its selected sub-graph, and Tasks counts the single tasks built from these steps (Section 3.1 ).
Task
Short name
Study
Kind
Codes
Checks
202509_001
hbn-sliding-sweep-inputs
1
input preparation
VASP
14
202509_002
hbn-sliding-potential-step
1
output analysis
VASP
4
202509_003
hbn-undulated-sliding-ratio
1
run and analyze
–
8
202509_004
bn-nanotube-inputs
1
input preparation
VASP
15
202509_005
bn-nanotube-flexoelectric-constant
1
output analysis
VASP
1
202509_006
hbn-flexo-sliding-decomposition-inputs
1
input preparation
VASP
14
Appendix
Table 10
Model
Harness
Steps
Context
Output
Effort
DeepSeek V4.1 Flash
MatClaw
100
800k
131,072
high
Gemini 3.8 Flash
MatClaw
500
256k
65,536
high
GPT-5.6 Sol
MatClaw
500
256k
128,000
high
GPT-5.6 Sol
Codex CLI 0.154.0
none
256k
–
high
Appendix
Table B1: Agent configuration. Steps is the step limit per single task, Context the context window, Output the maximum output tokens per call, and Effort the reasoning effort. For Codex CLI, none means that its native tool loop has no step cap, and the dash means its default output limit.
Simulation code
Version
Single tasks
Workflows
VASP
6.1.0; 6.4.1
no
yes
Quantum ESPRESSO
7.5
no
yes
ABACUS
3.10.1
no
yes
Wannier90
3.1.0
no
yes
LAMMPS
22 Jul 2025, update 4
yes
yes
phono3py, with phonopy
4.4.0
yes
yes
Appendix
Table B2: Software in the container. Single tasks and Workflows state whether the agent can run each simulation code: yes, no, or source when only a read-only source tree is staged. Where it cannot, the agent has the code’s manual. The Python packages are shared by every run.
Workflow
Study
Step
Description
Single task
C01
8
a
Build and relax the WS 2 /WSe 2 moiré
202602_015 to 016
b
Compute the piezoelectric bound charge
202602_021
c
Compute the screened electron potential
202602_022
C02
4
a
Classify Se-pair substitutions by space group
202602_025
b
Build a tight-binding model per family
–
c
Compute the spin splittings
202602_028 to 029
Appendix
Table 13
Figure C1: Complete results. Each square is one run. DeepSeek and Gemini are DeepSeek V4.1 Flash and Gemini 3.8 Flash; Sol and Codex CLI are GPT-5.6 Sol in the two harnesses of Table B . (a) Columns are the 94 single tasks of Table A , labelled by task number. SR rows are blank for the 42 tasks whose prompt has no reduced form.
Relaxation without a van der Waals correction ( IVDW unset)
A dispersion correction for the van der Waals-bonded bilayer
all four
202603_011, CrN Wannier functions to exchange constants, inputs
VASP, Wannier90, TB2J
Spin-channel files named wannier90.up and wannier90.dn
The names VASP writes, wannier90.1 and wannier90.2
DeepSeek, Sol, Codex CLI
202603_014, CrN Monte Carlo deck
VAMPIRE
No unit-cell-category in the material file, so only one Cr sublattice is built
Each sublattice assigned its own category
DeepSeek, Sol, Codex CLI
202601_002, V 2 CoAl relaxation inputs
VASP
LASPH not set
LASPH = .TRUE. for 3d elements
DeepSeek, Gemini
202603_009, CrN DFT+U structure inputs
VASP
LASPH not set
LASPH = .TRUE. for 3d elements with DFT+U
DeepSeek, Gemini
Appendix
Table D1: Errors shared across agents on single tasks under full guidance. Task numbers refer to Table A . Agents names those that fail the task with the error shown; an agent not named passes the task or fails it in another way.
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Suzhou Laboratory, Suzhou, China · State Key Laboratory for General Artificial Intelligence, BIGAI, Beijing, China +3