cs.AIJul 25, 2026

Stress-testing large language model agents in a robotic chemistry laboratory

Authors: Lulu GuoYingkai SunXiaobo LiLuyao GeZiming WangHaitao ZhengJingyu LiHuijuan Zhang+8 more

Organizations: State Key Laboratory of Precision and Intelligent Chemistry, Hefei National Research Center for Physical Sciences at the Microscale, School of Chemistry and Materials Science, University of Science and Technology of China, Hefei, China · Center for Scientific Intelligence Innovation, Hefei, China · School of Chemistry, School of Computer Science, University of Birmingham, Birmingham, UK

Abstract

AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.

Explore similar work

CardsList
  1. onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

    Aug 3, 2026Brandon Wang, Andrei S. Tyrin, Daniil A. BoikoVast Chemical Space