Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.
Figures & tables
Figure 1 : Stage dependence in generative VLA action formation. (A) Conceptual illustration of the evolving conditioning focus across denoising stages, from joint instruction–observation conditioning toward stronger visual alignment. (B) Relative attention to instruction and image tokens across denoising stages in π0 on LIBERO, showing an increasing emphasis on visual observations during later refinement. (C) Task success of partially denoised actions generated by GR00T on LIBERO, showing that substantial task competence emerges before denoising is complete.
Figure 2 : Overall framework of StairVLA. The high-level VLA performs early denoising over a long-horizon trajectory and caches the partially denoised trajectory for reuse across multiple steps. For each local chunk, the refiner uses the corresponding segment and the latest observation to complete refinement. The top, middle, and bottom sequences denote the initial noise, partially denoised trajectory, and executable action chunk, respectively.
Figure 3 : Overview of the low-level refiner. (A) The refiner uses the current observation, compressed high-level VLM tokens, and partially denoised action context to refine the query chunk. (B) The refiner adopts a GR00T-style DiT with interleaved cross-attention and self-attention for conditional fusion and action modeling.
Figure 4 : Refiner training path construction. After early denoising by the high-level VLA, the resulting partially denoised action is used to construct the refinement path toward the ground-truth action, from which both refiner inputs and targets are generated.
Method
Spatial
Object
Goal
Long
Avg.
Latency
Representative VLA Policies
OpenVLA ( Kim et al., 2024 )
84.7
88.4
79.2
53.7
76.5
–
π0 ( Black et al., 2024 )
96.8
98.8
95.8
85.2
94.1
–
GR00T ( Bjorck et al., 2025 )
94.4
97.6
93.0
90.6
93.9
–
π0.5 ( Physical Intelligence et al., 2025 )
98.8
98.2
98.0
92.4
96.9
–
Hierarchical / Coarse-to-Fine VLA Policies
Table 1 : Comparison on the LIBERO benchmark. We report success rates (%) on the four LIBERO suites, the average success rate (%), and reported inference latency (ms/chunk) when available. Chunk sizes may vary across methods. Best and second-best results are shown in bold and underlined , respectively. ∗ Uses additional subtask-level supervision beyond standard action demonstrations and is excluded from best/second-best ranking. ‡ StreamVLA reports average wall-clock latency under its gated reasoning schedule. § CF-VLA reports action-sampling latency excluding visual-language prefix encoding and KV-cache construction.
Method
Camera
Robot
Language
Light
Background
Noise
Layout
Avg.
Zero-Shot Transfer
OpenVLA ( Kim et al., 2024 )
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
π0 -FAST ( Pertsch et al., 2025 )
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
OpenVLA-OFT ( Kim et al., 2025 )
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
GR00T ( Bjorck et al., 2025 )
34.6
50.8
85.1
86.5
86.5
63.0
73.6
66.8
Libra-VLA ( Wei et al., 2026 )
68.9
48.8
92.7
97.9
93.4
86.3
77.5
79.5
Table 2: Results on LIBERO-Plus under zero-shot transfer and supervised fine-tuning.
Figure 7
Setting
Refiner
Action Context
Spatial
Object
Goal
Long
Avg.
Δ
StarVLA-GR00T ( StarVLA Community, 2026 )
–
–
97.8
98.8
97.4
92.0
96.5
–
Ours w/o refiner (Top-only)
✗
–
96.8
89.4
96.8
86.4
92.4
-4.1
Ours w/o Action Context
✓
✗
99.2
98.6
95.6
95.4
97.2
+0.7
Ours
✓
✓
98.0
99.6
97.8
95.8
97.8
+1.3
Table 3: Ablations on LIBERO.
Figure 9Figure 10
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11 : Denoising-stage attention analysis for GR00T and π0 . We record action-query cross-attention to different conditioning-token groups over 100 dense denoising stages. Both models exhibit stage-dependent attention patterns, with the changes being more pronounced in the first action-head layer.
Figure 12 : Representative first-layer cross-attention maps of π0 and GR00T at different denoising stages. For π0 , attention to image tokens becomes stronger toward later stages, while task-instruction attention does not exhibit a clear monotonic trend. For GR00T, late-stage cross-attention does not simply shift toward raw image or task-instruction tokens, indicating that different VLA action heads may realize stage-dependent conditioning through different token pathways.
Figure 13 : First-layer self-attention allocation in GR00T across denoising stages. Attention to action keys increases toward later stages, while attention to state/future keys decreases.
Method
M
Chunk Size
Spatial
Object
Goal
Long
Avg.
Δ Avg.
Latency
GR00T-style backbone
StarVLA-GR00T
–
8
97.8
98.8
97.4
92.0
96.5
–
115.0
Ours (GR00T-base)
2
5
98.4
98.8
97.6
94.8
97.40
+0.90
68.7
Ours (GR00T-base)
3
5
99.2
98.8
97.8
94.0
97.45
+0.95
52.3
Ours (GR00T-base)
4
5
98.0
99.6
97.8
95.8
97.80
+1.30
44.2
Ours (GR00T-base)
5
5
98.4
98.8
97.4
93.6
97.05
+0.55
39.3
Appendix
Table 4: Complete LIBERO results across GR00T-style and π -style VLA backbones. M denotes the number of 5-action low-level chunks executed between two high-level policy updates. The hierarchical GR00T- and π -based models use high-level action horizons of H=32 and H=20 , respectively. Latency is reported per native returned action chunk. The non-hierarchical StarVLA baselines use 8-action chunks, whereas our low-level refiner outputs 5-action chunks.
Figure 14 : Effect of long-horizon high-level trajectory reuse. We report the average LIBERO success rate as the cached high-level trajectory is reused for increasingly long execution horizons. Both curves are shown from 25 executed actions onward, corresponding to the shared evaluation range of the H=48 and H=64 models.
Figure 16
Task ID
Benchmark
Task Description
Retained Demonstrations
1
Fruit25
Pick up the apple and place it into the red basket.
100
2
Fruit25
Pick up the apple and place it onto the metal tray.
100
3
Fruit25
Pick up the orange and place it onto the metal tray.
100
4
Fruit25
Pick up the orange and place it into the orange basket.
96
5
Fruit25
Pick up the lemon and place it onto the metal tray.
98
6
Fruit25
Pick up the lemon and place it into the yellow basket.
99
Appendix
Table 5: Task definitions and retained demonstration counts for the real-world datasets.
ID
Task
StarVLA-GR00T
StarVLA- π
Ours
Fruit25
1
Apple → Red basket
7/10 (70%)
6/10 (60%)
10/10 (100%)
2
Apple → Metal tray
10/10 (100%)
9/10 (90%)
9/10 (90%)
3
Orange → Metal tray
9/10 (90%)
9/10 (90%)
9/10 (90%)
4
Orange → Orange basket
5/10 (50%)
8/10 (80%)
7/10 (70%)
5
Lemon → Metal tray
10/10 (100%)
10/10 (100%)
10/10 (100%)
Appendix
Table 6: Per-task success results on the real-world benchmarks. Fruit25 tasks are evaluated with either 5 or 10 trials, while PushBlock is evaluated with 10 trials per method.
Figure 17 : Examples of relatively easy PushBlock initial configurations, where the block is approximately aligned with the target region and can be pushed along a near-straight trajectory.
Method
T-SPARC ↑
R-SPARC ↑
T-Jerk ↓
R-Jerk ↓
Stop (%) ↓
StarVLA-GR00T
-5.480 [-5.817, -5.127]
-10.499 [-11.932, -9.440]
9.599 [9.408, 9.788]
10.026 [9.833, 10.216]
12.537 [10.575, 15.102]
StarVLA- π
-6.799 [-7.266, -6.233]
-12.134 [-13.370, -11.135]
9.939 [9.681, 10.212]
10.355 [10.091, 10.599]
26.066 [23.566, 28.767]
Ours
-5.151 [-5.398, -4.887]
-9.992 [-11.099, -8.959]
9.484 [9.300, 9.644]
9.909 [9.733, 10.091]
11.274 [8.974, 13.607]
Appendix
Table 7: Motion smoothness and execution continuity on Fruit25 trials. Values are reported as median [Q1, Q3]. Higher SPARC and lower jerk and stop ratio indicate smoother and more continuous execution.