cs.ROOct 6, 2026

StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models

Authors: Shangyuan Yuan, Xinda Qi, Yujiang Pu, Wenliang Guo, Xiaobo Tan

Organizations: Michigan State University · Ant Group

Abstract

Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

    Apr 27, 2026Fan Du, Feng Yan, Jianxiong Wu +8Robotic ControlEfficient VLA Model Inference

  2. Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Jul 14, 2026Yuzhou Wu, Yuxin Zheng, Muchun Niu +6Efficient VLA Model InferenceEfficient VLA Models

  3. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Robotic ControlEfficient VLA Model Inference