cs.ROMar 18, 2026

ExecVLA: Following Fine-Grained Execution Constraints in Vision-Language-Action Models with Bi-Level Action Representation

Authors: Gaoge Han, Zhengqing Gao, Ziwen Li, Jiaxin Huang, Shaoli Huang, Fakhri Karray, Mingming Gong, Tongliang Liu

Organizations: MBZUAI.

Abstract

We study fine-grained execution-constraint following in vision-language-action (VLA) models. Given an invariant task goal, the policy must follow instruction-specified execution constraints, including interaction targets, motion patterns, spatial relations, and terminal configurations. This setting exposes a limitation of goal-oriented VLAs: trajectories that complete the same task are not interchangeable when the instruction specifies how the task must be executed. We propose ExecVLA, a framework that separates a goal-oriented component from an execution-specific component through a bi-level action representation and supervised bi-level reasoning tokens. We further introduce explicit goal-invariance and execution-predictability objectives so that the goal-level representation remains stable across executions of the same goal, while the execution-level representation retains the constraints that distinguish those executions. We construct execution-constraint-following datasets in simulation and on a Realman-75 robot, with goal and fine-grained reasoning annotations. Experiments on LIBERO and the real robot show improved goal completion and, more importantly, substantially more reliable adherence to instruction-specified execution constraints.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies

    May 26, 2026Xintong Hu, Xuhong Huang, Jinyu Zhang +11Language-Conditioned Robot ManipulationVision-Language-Action Models

  2. Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

    Aug 5, 2026Xingyu Ding, Yuzhong Zhao, Yang Wu +4Robotic ManipulationVision-Language-Action Models

  3. Dynamic Execution Commitment of Vision-Language-Action Models

    May 12, 2026Feng Chen, Xianghui Wang, Yuxuan Chen +4VLM RobustnessEfficient VLM Inference