cs.CVSep 27, 2026

EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning

Authors: Jinchang Zhang, Yingda Tao, Jiakai Lin, Guoyu Lu

Organizations: Intelligent Vision and Sensing (IVS) Lab Indiana University Bloomington Bloomington, IN, USA

Abstract

Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state to reason about relations, constraints, and the consequences of design changes. We introduce \textsc{EngIntervene}, a benchmark for this capability. It contains 3,229 questions across seven engineering domains and organizes evaluation into four levels: state grounding (T1), relational and mechanistic reasoning (T2), constraint-aware diagnosis (T3), and intervention reasoning (T4), which asks whether a proposed modification achieves its target while preserving required constraints. The tasks instantiate a unified engineering state representation spanning objects, relations, constraints, and design objectives, and T2--T4 are scored against structured reference answers with atomic criteria. Across open- and closed-weight multimodal models, stronger grounding does not reliably translate into better diagnosis or intervention, and the best open-weight model trails the best closed model by 14.7 percentage points on the T2--T4 average. Removing or shuffling visual evidence consistently degrades performance, while benchmark-specific supervised fine-tuning improves T1 but not T2--T4. T4 further exposes a large gap between satisfying individual revision criteria and producing a fully valid intervention. Engineering reasoning thus requires not only recovering the current state, but also reliably using it to reason about constraints and post-intervention consequences. Code and benchmark artifacts are available at https://github.com/changcv2021/EngIntervene

Figures & tables

Explore similar work

CardsList
  1. Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

    Jun 9, 2026Syed Wasiq, Syed Mohamad Tawseeq, Yashwant Pravinrao Bangde +1Multimodal Reasoning BenchmarksMultimodal Reasoning

  2. MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

    Aug 10, 2026Chenxu Du, Kang An, Tengyue Wang +8Multimodal Large Language ModelsMultimodal Reasoning

  3. See2Think: Do Multimodal Models Really Use Intermediate Visual States?

    Jul 29, 2026Siyu Yan, Zhuoran Yan, Haiying Xu +10Large Multimodal ModelsMultimodal Model