cs.AISep 29, 2026

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

Authors: Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, +14 more

Organizations: Tsinghua University · Zhiman Inc. · Chongqing University · UCAS · Shandong University · BUPT · Fudan University · Henan Polytechnic University · Xi’an Jiaotong University · Peking University · Zhejiang University

Abstract

Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

    May 19, 2026Gioele Molinari, Florian Felten, Soheyl Massoudi +1Engineering

  2. DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness

    Jul 30, 2026Debin Meng, Jiaming Yang, Zefang Zong +4Data Science AgentsMle-Bench Lite

  3. SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

    Sep 8, 2026Pujun Zheng, Zixin Shang, Shufan Jiang +5Swe-Bench VerifiedRepository-Level Code Understanding