CompMat-Bench: Benchmarking AI Agents for Computational Materials Science
Organizations: Department of Materials Science and NanoEngineering Rice University, Houston, TX 77005, USA
Abstract
Evaluating AI agents on scientific research tasks is constrained by the time and resources required for the underlying experiments or calculations. In computational materials research, repeating the same expensive simulations across agents and trials can make evaluation impractical. We introduce CompMat-Bench, a benchmark of 94 tasks derived from recently published computational materials studies, each asking agents to complete a step toward achieving the study's scientific goal. We reproduce the research steps in advance and assess agents on preparing inputs and analyzing outputs for expensive simulations, so expensive simulations can be avoided during evaluation. The reproduced inputs and results serve as ground truth for grading agents with fixed rules, without an LLM judge. The benchmark supports four evaluation conditions: single tasks and workflows composed of related tasks, each with full or reduced methodological guidance. With full guidance on single tasks, agents based on three LLMs demonstrate the ability to complete individual materials research steps, with pass rates of 66.0-90.4% across 94 tasks. Both longer workflows and reduced guidance can limit agent performance, but in different ways for different agents: they lower the pass rates of the weaker agents, whereas the strongest agent falls only when a long workflow is combined with reduced guidance. Failure analysis attributes most failures to scientific errors rather than to errors in software usage. CompMat-Bench provides a basis for comparing agents on the steps of real materials research and for analyzing agent failure modes.
Figures & tables
| Published | Expensive | Excluded from | |||
|---|---|---|---|---|---|
| Benchmark | Domain | studies | computation | evaluation | Grading |
| FIRE-Bench ( Wang et al., 2026b ) | ML | ✓ | ✗ | – | LLM judge |
| ReplicationBench ( Ye et al., 2025 ) | astrophysics | ✓ | (✓) | (✓) | fixed rules |
| PaperBench ( Starace et al., 2025 ) | ML | ✓ | (✓) | ✗ | LLM judge |
| AutoMat ( Huang et al., 2026a ) | materials | ✓ | ✓ | (✓) | LLM judge |
| SimulCost ( Cao et al., 2026 ) | physics sim. | ✗ | (✓) | ✗ | fixed rules |
| Pass rate (%) | ||||||
| Prepare | Analyze | Run and | All | API cost | ||
| Model | Harness | inputs | outputs | analyze | tasks | (USD) |
| DeepSeek V4.1 Flash | MatClaw | 36.7 | 79.3 | 80.0 | 66.0 (62/94) | 3.86 |
| Gemini 3.8 Flash | MatClaw | 70.0 | 96.6 | 100.0 | 89.4 (84/94) | 106.43 |
| GPT-5.6 Sol | MatClaw | 76.7 | 96.6 | 97.1 | 90.4 (85/94) | 27.00 |
| Tasks by passing runs | |||||||
| Model | Pass rate (%) | 5/5 | 4/5 | 3/5 | 2/5 | 1/5 | 0/5 |
| DeepSeek V4.1 Flash | 62.0 | 7 | 3 | 2 | 3 | 3 | 2 |
| Gemini 3.8 Flash | 88.0 | 16 | 1 | 1 | 0 | 1 | 1 |
| GPT-5.6 Sol | 93.0 | 17 | 2 | 0 | 0 | 0 | 1 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| No. | Study | Journal | Month | Methods | Graph steps | Selected | Tasks |
| 1 | ( Li et al., 2025 ) | Phys. Rev. B | 2025-09 | DFT, Structure | 10 | 4 | 7 |
| 2 | ( Jiang et al., 2026 ) | Nano Lett. | 2026-01 | Model Hamiltonians, DFT | 7 | 5 | 5 |
| 3 | ( Xie et al., 2026 ) | Phys. Rev. B | 2026-01 | DFT, Structure, Spin models | 10 | 5 | 8 |
| 4 | ( Devaraj et al., 2026 ) | Phys. Rev. B | 2026-02 | DFT, Model Hamiltonians, Structure | 12 | 5 | 7 |
| 5 | ( Hofer et al., 2026 ) | Nano Lett. | 2026-02 | Model Hamiltonians | 6 | 3 | 4 |
| 6 | ( Parlak et al., 2026 ) | Phys. Rev. B | 2026-02 | Phonons, DFT | 10 | 3 | 6 |
| Task | Short name | Study | Kind | Codes | Checks |
|---|---|---|---|---|---|
| 202509_001 | hbn-sliding-sweep-inputs | 1 | input preparation | VASP | 14 |
| 202509_002 | hbn-sliding-potential-step | 1 | output analysis | VASP | 4 |
| 202509_003 | hbn-undulated-sliding-ratio | 1 | run and analyze | – | 8 |
| 202509_004 | bn-nanotube-inputs | 1 | input preparation | VASP | 15 |
| 202509_005 | bn-nanotube-flexoelectric-constant | 1 | output analysis | VASP | 1 |
| 202509_006 | hbn-flexo-sliding-decomposition-inputs | 1 | input preparation | VASP | 14 |
| Model | Harness | Steps | Context | Output | Effort |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | MatClaw | 100 | 800k | 131,072 | high |
| Gemini 3.8 Flash | MatClaw | 500 | 256k | 65,536 | high |
| GPT-5.6 Sol | MatClaw | 500 | 256k | 128,000 | high |
| GPT-5.6 Sol | Codex CLI 0.154.0 | none | 256k | – | high |
| Simulation code | Version | Single tasks | Workflows |
|---|---|---|---|
| VASP | 6.1.0; 6.4.1 | no | yes |
| Quantum ESPRESSO | 7.5 | no | yes |
| ABACUS | 3.10.1 | no | yes |
| Wannier90 | 3.1.0 | no | yes |
| LAMMPS | 22 Jul 2025, update 4 | yes | yes |
| phono3py, with phonopy | 4.4.0 | yes | yes |
| Workflow | Study | Step | Description | Single task |
|---|---|---|---|---|
| C01 | 8 | a | Build and relax the WS 2 /WSe 2 moiré | 202602_015 to 016 |
| b | Compute the piezoelectric bound charge | 202602_021 | ||
| c | Compute the screened electron potential | 202602_022 | ||
| C02 | 4 | a | Classify Se-pair substitutions by space group | 202602_025 |
| b | Build a tight-binding model per family | – | ||
| c | Compute the spin splittings | 202602_028 to 029 |
| Task | Code | What the agents write | What the task needs | Agents |
|---|---|---|---|---|
| 202602_015, WS 2 /WSe 2 moiré bilayer, LAMMPS inputs | LAMMPS | Interlayer potential kolmogorov/crespi/z with two element names | One element name per atom type, six for this cell | all four |
| 202602_017, WS 2 /WSe 2 bilayer, band-structure inputs | VASP | Relaxation without a van der Waals correction ( IVDW unset) | A dispersion correction for the van der Waals-bonded bilayer | all four |
| 202603_011, CrN Wannier functions to exchange constants, inputs | VASP, Wannier90, TB2J | Spin-channel files named wannier90.up and wannier90.dn | The names VASP writes, wannier90.1 and wannier90.2 | DeepSeek, Sol, Codex CLI |
| 202603_014, CrN Monte Carlo deck | VAMPIRE | No unit-cell-category in the material file, so only one Cr sublattice is built | Each sublattice assigned its own category | DeepSeek, Sol, Codex CLI |
| 202601_002, V 2 CoAl relaxation inputs | VASP | LASPH not set | LASPH = .TRUE. for 3d elements | DeepSeek, Gemini |
| 202603_009, CrN DFT+U structure inputs | VASP | LASPH not set | LASPH = .TRUE. for 3d elements with DFT+U | DeepSeek, Gemini |