Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.
Figures & tables
Figure 1 : Stage dependence in generative VLA action formation. (A) Conceptual illustration of the evolving conditioning focus across denoising stages, from joint instruction–observation conditioning toward stronger visual alignment. (B) Relative attention to instruction and image tokens across denoising stages in π0 on LIBERO, showing an increasing emphasis on visual observations during later refinement. (C) Task success of partially denoised actions generated by GR00T on LIBERO, showing that substantial task competence emerges before denoising is complete.
Figure 2 : Overall framework of StairVLA. The high-level VLA performs early denoising over a long-horizon trajectory and caches the partially denoised trajectory for reuse across multiple steps. For each local chunk, the refiner uses the corresponding segment and the latest observation to complete refinement. The top, middle, and bottom sequences denote the initial noise, partially denoised trajectory, and executable action chunk, respectively.
Figure 3 : Overview of the low-level refiner. (A) The refiner uses the current observation, compressed high-level VLM tokens, and partially denoised action context to refine the query chunk. (B) The refiner adopts a GR00T-style DiT with interleaved cross-attention and self-attention for conditional fusion and action modeling.
Figure 4 : Refiner training path construction. After early denoising by the high-level VLA, the resulting partially denoised action is used to construct the refinement path toward the ground-truth action, from which both refiner inputs and targets are generated.
Method
Spatial
Object
Goal
Long
Avg.
Latency
Representative VLA Policies
OpenVLA ( Kim et al., 2024 )
84.7
88.4
79.2
53.7
76.5
–
π0 ( Black et al., 2024 )
96.8
98.8
95.8
85.2
94.1
–
GR00T ( Bjorck et al., 2025 )
94.4
97.6
93.0
90.6
93.9
–
π0.5 ( Physical Intelligence et al., 2025 )
98.8
98.2
98.0
92.4
96.9
–
Hierarchical / Coarse-to-Fine VLA Policies
Table 1 : Comparison on the LIBERO benchmark. We report success rates (%) on the four LIBERO suites, the average success rate (%), and reported inference latency (ms/chunk) when available. Chunk sizes may vary across methods. Best and second-best results are shown in bold and underlined , respectively. ∗ Uses additional subtask-level supervision beyond standard action demonstrations and is excluded from best/second-best ranking. ‡ StreamVLA reports average wall-clock latency under its gated reasoning schedule. § CF-VLA reports action-sampling latency excluding visual-language prefix encoding and KV-cache construction.
Method
Camera
Robot
Language
Light
Background
Noise
Layout
Avg.
Zero-Shot Transfer
OpenVLA ( Kim et al., 2024 )
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
π0 -FAST ( Pertsch et al., 2025 )
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
OpenVLA-OFT ( Kim et al., 2025 )
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
GR00T ( Bjorck et al., 2025 )
34.6
50.8
85.1
86.5
86.5
63.0
73.6
66.8
Libra-VLA ( Wei et al., 2026 )
68.9
48.8
92.7
97.9
93.4
86.3
77.5
79.5
Table 2: Results on LIBERO-Plus under zero-shot transfer and supervised fine-tuning.
Figure 7
Setting
Refiner
Action Context
Spatial
Object
Goal
Long
Avg.
Δ
StarVLA-GR00T ( StarVLA Community, 2026 )
–
–
97.8
98.8
97.4
92.0
96.5
–
Ours w/o refiner (Top-only)
✗
–
96.8
89.4
96.8
86.4
92.4
-4.1
Ours w/o Action Context
✓
✗
99.2
98.6
95.6
95.4
97.2
+0.7
Ours
✓
✓
98.0
99.6
97.8
95.8
97.8
+1.3
Table 3: Ablations on LIBERO.
Figure 9Figure 10
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11 : Denoising-stage attention analysis for GR00T and π0 . We record action-query cross-attention to different conditioning-token groups over 100 dense denoising stages. Both models exhibit stage-dependent attention patterns, with the changes being more pronounced in the first action-head layer.
Figure 12 : Representative first-layer cross-attention maps of π0 and GR00T at different denoising stages. For π0 , attention to image tokens becomes stronger toward later stages, while task-instruction attention does not exhibit a clear monotonic trend. For GR00T, late-stage cross-attention does not simply shift toward raw image or task-instruction tokens, indicating that different VLA action heads may realize stage-dependent conditioning through different token pathways.
Figure 13 : First-layer self-attention allocation in GR00T across denoising stages. Attention to action keys increases toward later stages, while attention to state/future keys decreases.
Method
M
Chunk Size
Spatial
Object
Goal
Long
Avg.
Δ Avg.
Latency
GR00T-style backbone
StarVLA-GR00T
–
8
97.8
98.8
97.4
92.0
96.5
–
115.0
Ours (GR00T-base)
2
5
98.4
98.8
97.6
94.8
97.40
+0.90
68.7
Ours (GR00T-base)
3
5
99.2
98.8
97.8
94.0
97.45
+0.95
52.3
Ours (GR00T-base)
4
5
98.0
99.6
97.8
95.8
97.80
+1.30
44.2
Ours (GR00T-base)
5
5
98.4
98.8
97.4
93.6
97.05
+0.55
39.3
Appendix
Table 4: Complete LIBERO results across GR00T-style and π -style VLA backbones. M denotes the number of 5-action low-level chunks executed between two high-level policy updates. The hierarchical GR00T- and π -based models use high-level action horizons of H=32 and H=20 , respectively. Latency is reported per native returned action chunk. The non-hierarchical StarVLA baselines use 8-action chunks, whereas our low-level refiner outputs 5-action chunks.
Figure 14 : Effect of long-horizon high-level trajectory reuse. We report the average LIBERO success rate as the cached high-level trajectory is reused for increasingly long execution horizons. Both curves are shown from 25 executed actions onward, corresponding to the shared evaluation range of the H=48 and H=64 models.
Figure 16
Task ID
Benchmark
Task Description
Retained Demonstrations
1
Fruit25
Pick up the apple and place it into the red basket.
100
2
Fruit25
Pick up the apple and place it onto the metal tray.
100
3
Fruit25
Pick up the orange and place it onto the metal tray.
100
4
Fruit25
Pick up the orange and place it into the orange basket.
96
5
Fruit25
Pick up the lemon and place it onto the metal tray.
98
6
Fruit25
Pick up the lemon and place it into the yellow basket.
99
Appendix
Table 5: Task definitions and retained demonstration counts for the real-world datasets.
ID
Task
StarVLA-GR00T
StarVLA- π
Ours
Fruit25
1
Apple → Red basket
7/10 (70%)
6/10 (60%)
10/10 (100%)
2
Apple → Metal tray
10/10 (100%)
9/10 (90%)
9/10 (90%)
3
Orange → Metal tray
9/10 (90%)
9/10 (90%)
9/10 (90%)
4
Orange → Orange basket
5/10 (50%)
8/10 (80%)
7/10 (70%)
5
Lemon → Metal tray
10/10 (100%)
10/10 (100%)
10/10 (100%)
Appendix
Table 6: Per-task success results on the real-world benchmarks. Fruit25 tasks are evaluated with either 5 or 10 trials, while PushBlock is evaluated with 10 trials per method.
Figure 17 : Examples of relatively easy PushBlock initial configurations, where the block is approximately aligned with the target region and can be pushed along a near-straight trajectory.
Method
T-SPARC ↑
R-SPARC ↑
T-Jerk ↓
R-Jerk ↓
Stop (%) ↓
StarVLA-GR00T
-5.480 [-5.817, -5.127]
-10.499 [-11.932, -9.440]
9.599 [9.408, 9.788]
10.026 [9.833, 10.216]
12.537 [10.575, 15.102]
StarVLA- π
-6.799 [-7.266, -6.233]
-12.134 [-13.370, -11.135]
9.939 [9.681, 10.212]
10.355 [10.091, 10.599]
26.066 [23.566, 28.767]
Ours
-5.151 [-5.398, -4.887]
-9.992 [-11.099, -8.959]
9.484 [9.300, 9.644]
9.909 [9.733, 10.091]
11.274 [8.974, 13.607]
Appendix
Table 7: Motion smoothness and execution continuity on Fruit25 trials. Values are reported as median [Q1, Q3]. Higher SPARC and lower jerk and stop ratio indicate smoother and more continuous execution.
Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to recover action structure from uninformative Gaussian noise, leading to a poor efficiency-quality trade-off under real-time constraints. We address this issue by rethinking the role of the starting point in generative action modeling. Instead of shortening the sampling trajectory, we propose CF-VLA, a coarse-to-fine two-stage formulation that restructures action generation into a coarse initialization step that constructs an action-aware starting point, followed by a single-step local refinement that corrects residual errors. Concretely, the coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed-time refinement from this initialization. To stabilize training, we introduce a stepwise strategy that first learns a controlled coarse predictor and then performs joint optimization. Experiments on CALVIN and LIBERO show that our method establishes a strong efficiency-performance frontier under low-NFE (Number of Function Evaluations) regimes: it consistently outperforms existing NFE=2 methods, matches or surpasses the NFE=10 π0.5 baseline on several metrics, reduces action sampling latency by 75.4%, and achieves the best average real-robot success rate of 83.0%, outperforming MIP by 19.5 points and π0.5 by 4.0 points. These results suggest that structured, coarse-to-fine generation enables both strong performance and efficient inference. Our code is available at https://github.com/EmbodiedAI-RoboTron/CF-VLA.
Fan Du, Feng Yan, Jianxiong Wu +8
1Southern University of Science and Technology · 2Xi’an Jiaotong University · 3United Nova Technology +1
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading to high inference latency and increased training cost, and (ii) insufficient utilization of generated actions resulting in potential performance loss. To address these issues, we develop a training framework to finetune VLA models for generating significantly fewer action tokens with high parallelism, effectively reducing inference latency and training cost. Furthermore, we introduce an inference optimization technique with a novel voting-based ensemble strategy to combine current and previous action predictions, improving the utilization of generated actions and overall performance. Our results demonstrate that we achieve superior performance compared with state-of-the-art VLA models, achieving significantly higher success rates and 39× faster inference than OpenVLA with 46 Hz throughput on edge platforms, demonstrating practical deployability. The code is available at https://github.com/LukeLIN-web/VOTE.
Juyi Lin, Amir Taherin, Arash Akbari +11
Northeastern University, Boston, USA · EmbodyX,San Mateo,USA