Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find that discarded actions stay close to their replanned versions as long as the action velocity remains smooth. Action Upcycling therefore extends the execution horizon up to the point where the velocity begins to fluctuate. Extensive experiments on simulated and real-world manipulation tasks show that Action Upcycling reduces policy calls by 1.2-1.7x with no loss in success rate, across multiple Vision-Language-Action Models (VLAs) and even a World Action Model (WAM). It applies to any chunked policy at negligible cost and is orthogonal to other policy acceleration methods such as few-step sampling and streaming action decoding, opening a new axis for policy acceleration.
Figures & tables
Figure 1: Discarded actions are worth keeping when motion is smooth. We compare the discarded actions of each chunk with their replanned versions on LIBERO for (a) SmolVLA, (b) GR00T N1.7, (c) π0.5 , and (d) FastWAM. ( left ) They are strongly correlated. ( right ) Their RMSE grows with the velocity fluctuation of the discarded actions. Action Upcycling exploits this property to decide how many discarded actions to reuse.
Figure 2: Overview of Action Upcycling. Top : A chunked policy predicts H actions per call, executes only the first h , and discards the rest (the tail ). Action Upcycling additionally executes the trustworthy part of the tail, raising the mean execution length from h to rh and reducing policy calls accordingly. Bottom : The accumulated velocity fluctuation ck stays small while the predicted motion is smooth and rises once it fluctuates. Action Upcycling executes tail actions while ck≤τ , where τ is determined from a pool C of signals ck , so that the mean execution length reaches rh .
Baseline Policy
+ Action Upcycling
Benchmark
Model
succ. ↑
calls / ep ↓
ms
s / ep ↓
succ. ↑
calls / ep ↓
ms
s / ep ↓
LIBERO
π0.5
96.9
32.4 (1 × )
136
4.41
97.9
21.9 (1.5 × )
137
3.00
SmolVLA
82.6
19.3 (1 × )
94
1.82
82.8
14.7 (1.3 × )
94
1.39
GR00T N1.7
96.1
22.6 (1 × )
115
2.59
96.6
15.0 (1.5 × )
115
1.72
FastWAM
97.4
15.7 (1 × )
84
1.32
97.6
13.4 (1.2 × )
84
1.12
LIBERO-Plus
π0.5
83.7
38.5 (1 × )
137
5.27
86.5
23.9 (1.6 × )
137
3.27
Table 1: Benefit of Action Upcycling. Results of four policies on three benchmarks. We report success rate (%), policy calls per episode, latency per policy call (ms), and inference time per episode (s). Action Upcycling consistently reduces policy calls by 1.2–1.7 × while matching or even improving the success rate. Better value in bold .
LIBERO
LIBERO-Plus
RoboTwin 2.0
Method
succ.
calls / ep
ms
s / ep
succ.
calls / ep
ms
s / ep
succ.
calls / ep
ms
s / ep
Baseline
96.9
32.4 (1 × )
136
4.41
83.7
38.5 (1 × )
137
5.27
59.7
40.9 (1 × )
152
6.22
+ AAC (CVPR 2026)
97.3
23.5 (1.38 × )
802
18.9
84.5
27.1 (1.42 × )
802
21.7
53.1
46.1 (0.89 × )
930
42.9
+ AutoHorizon (ECCV 2026)
97.4
22.3 (1.45 × )
136
3.03
85.5
24.5 (1.57 × )
136
3.33
56.8
28.0 (1.46 × )
152
4.26
+ Action Upcycling
97.9
21.9 (1.48 × )
137
3.00
86.5
23.9 (1.61 × )
137
3.27
60.5
29.8 (1.37 × )
152
4.53
Table 2: Comparison with adaptive execution horizon methods. Results of π0.5 on three benchmarks. We report success rate (%), policy calls per episode, latency per policy call (ms), and inference time per episode (s). Best in bold , second best underlined .
Setting
succ. ↑
calls / ep ↓
ms
s / ep ↓
speed-up ↑
Baseline ( N=10 )
96.9
32.4 (1 × )
136
4.41
1.0 ×
+ Action Upcycling
97.9
21.9 (1.48 × )
137
3.00
1.5 ×
N=5
96.9
32.4 (1 × )
97
3.14
1.4 ×
+ Action Upcycling
97.7
21.9 (1.48 × )
98
2.15
2.1 ×
N=2
96.8
32.0 (1 × )
69
2.21
2.0 ×
+ Action Upcycling
97.8
21.8 (1.47 × )
71
1.55
2.8 ×
Table 3: Action Upcycling stacks with denoising-step reduction. Results of π0.5 with N denoising steps on LIBERO. We report success rate (%), policy calls per episode, latency per policy call (ms), inference time per episode (s), and speed-up over the N=10 baseline. Better value at each N in bold .
Setting
succ. ↑
calls / ep ↓
ms
s / ep ↓
speed-up ↑
Baseline
96.90
32.4 (1 × )
136
4.41
1.0 ×
+ FlashVLA
98.65
31.4 (1.03 × )
32
1.01
4.4 ×
+ FlashVLA & Action Upcycling
98.70
21.5 (1.51 × )
31
0.66
6.7 ×
Table 4: Action Upcycling stacks with streaming action decoding. Results of π0.5 with FlashVLA on LIBERO. FlashVLA denoises a rolling buffer of chunks at different noise levels for faster decoding. We report success rate (%), policy calls per episode, latency per policy call (ms), inference time per episode (s), and speed-up over π0.5 . Best in bold .
Baseline
Fixed-length
Adaptive (ours)
Model
succ. ↑
hˉ
succ. ↑
hˉ
succ. ↑
hˉ
π0.5
96.90
5.0
97.05
7.0
97.90
7.3
SmolVLA
82.60
10.0
81.55
13.0
82.80
13.3
GR00T N1.7
96.05
8.0
95.40
12.0
96.55
11.9
FastWAM
97.35
10.0
96.45
12.0
97.55
11.9
Table 5: Adaptive vs. fixed-length execution. Results of four policies on LIBERO. Fixed-length execution uses a length close to the mean execution length hˉ of Action Upcycling. We report success rate (%) and hˉ . Better value in bold .
Figure 3: Distribution of execution lengths. The execution lengths selected by Action Upcycling spread over a range rather than concentrating on a single value.
Success rate ↑
Calls / ep ↓
Time / ep (s) ↓
Task
Baseline
Action Upcycling
Baseline
Action Upcycling
Baseline
Action Upcycling
Soccer ball → plate (layout 1)
10/10
10/10
26.9
26.8
14.4
17.4
Soccer ball → plate (layout 2)
10/10
10/10
36.5
20.7
19.0
12.9
Basketball → basket
10/10
10/10
31.3
24.9
16.5
15.8
Baseball → basket
10/10
10/10
32.1
25.0
17.1
16.1
Drawer: black ball → drawer
8/10
10/10
73.8
49.2
37.3
30.9
Table 6: Quantitative real-world results with π0.5 . Each task is evaluated over 10 episodes. We report success rate, policy calls per episode, and time per episode (s). Better value in bold .
Success rate ↑
Calls / ep ↓
Time / ep (s) ↓
Task
Baseline
Action Upcycling
Baseline
Action Upcycling
Baseline
Action Upcycling
Soccer ball → plate (layout 1)
10/10
10/10
56.8
50.3
17.8
17.7
Soccer ball → plate (layout 2)
8/10
9/10
102.2
81.7
31.5
28.8
Basketball → basket
8/10
8/10
95.6
69.9
29.6
24.3
Baseball → basket
8/10
8/10
107.9
82.3
33.1
28.9
Total / mean
34/40
35/40
90.6
71.0
28.0
24.9
Table 7: Quantitative real-world results with GR00T N1.6. Each task is evaluated over 10 episodes. We report success rate, policy calls per episode, and time per episode (s). Better value in bold .
Figure 4: Qualitative real-world results. A drawer task with π0.5 (top) and a pick-and-place task with GR00T N1.6 (bottom), each shown with the baseline policy and with Action Upcycling. Red and green boxes mark failure and success, respectively. In both tasks, the baseline policy times out before completion, while the same policy with Action Upcycling succeeds. Videos are available on the project page.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Model
LIBERO
LIBERO-Plus
RoboTwin 2.0
π0.5
5 / 10
5 / 10
10 / 32
SmolVLA
10 / 50
10 / 50
10 / 50
GR00T N1.7
8 / 16
8 / 16
—
FastWAM
10 / 32
10 / 32
24 / 32
Appendix
Table 8: Horizon settings. Execution horizon h and prediction horizon H of each model and benchmark.
Figure 5: Real-robot setup. Two YAM follower arms with wrist-mounted Intel RealSense D405 cameras, a top-mounted Intel RealSense D435, and two YAM leader arms used for teleoperated data collection. All experiments use only the right follower arm.
All chunks
2 % of chunks
Model
τ
hˉ
τ
hˉ
π0.5
0.15
7.3
0.15
7.3
SmolVLA
0.65
13.3
0.64
13.2
GR00T N1.7
0.22
11.9
0.22
11.9
FastWAM
0.09
11.9
0.09
11.9
Appendix
Table 9: A small pool suffices for selecting τ . Selected τ and resulting mean execution length hˉ on LIBERO, using pools built from all chunks and from a random 2 % of them.
Method
succ. ↑
calls / ep ↓
ms / call ↓
s / ep ↓
Baseline
96.9
32.4 (1 × )
136
4.41
+ Action Upcycling (offline pool C )
97.9
21.9 (1.48 × )
137
3.00
+ Action Upcycling (online pool C )
97.7
21.0 (1.54 × )
150
3.15
Appendix
Table 10: Offline vs. online pool construction. Results of π0.5 on LIBERO. Best in bold .
Figure 6: Effect of the upcycling ratio. Success rate (%) of each policy on LIBERO as r increases from the default setting ( r=1 ).
Figure 7: Additional qualitative real-world results with π0.5 . Four drawer tasks, each shown with the baseline policy (top) and with Action Upcycling (bottom). Red and green boxes mark failure and success, respectively. In two tasks, the baseline policy times out before completion, while the same policy with Action Upcycling succeeds. In the other two tasks, both succeed, but the policy with Action Upcycling completes the task faster.
Figure 8: Additional qualitative real-world results with GR00T N1.6. Four pick-and-place tasks, each shown with the baseline policy (top) and with Action Upcycling (bottom). Red and green boxes mark failure and success, respectively. In two tasks, the baseline policy times out before completion, while the same policy with Action Upcycling succeeds. In the other two tasks, both succeed, but the policy with Action Upcycling completes the task faster.
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, 3.62× the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
Guoheng Sun, Chen Chen, Jin Wang +2
University of Maryland, College Park · Independent Researcher · Oxford Robotics Institute, University of Oxford
Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execution efficiency. While existing efforts predominantly focus on compute-centric efficiency to reduce per-step inference latency, the intrinsic \textbf{policy efficiency} of these models remains largely unexplored. Policy efficiency is fundamentally affected by two factors, namely the effective executable length of predicted action chunks and the total physical steps required to complete a task. These two factors jointly determine the total number of forward inference calls during execution. We observe that current VLA policies struggle with planning unreliability and action redundancy, suffering from severe prediction degradation at the tail of action chunks and tending to generate unnecessarily redundant physical steps. To address this, we propose \textbf{PolicyTrim}, a reinforcement learning-based post-training framework that extends the reliable action chunk length and reduces redundant physical steps. For reliable chunk extension, we employ a dynamic exploration strategy that explicitly rewards the successful completion of longer executable lengths, progressively pushing the trustworthy prediction horizon to its empirical limit. For step efficiency, we design a redundancy-aware reward that directly favors successful task completions with fewer steps while penalizing unreproducible shortcuts, effectively eliminating redundant physical actions. Extensive experiments across three benchmarks and three VLA models demonstrate that PolicyTrim improves action chunk utilization by 3× and reduces physical execution steps by 51.4%. Ultimately, our framework delivers up to a 5.83× end-to-end deployment speedup without compromising task success rates.
Xianghui Wang, Feng Chen, Wenbo Zhang +4
1Sichuan University · 2Adelaide University · 3Beijing Institute of Technology
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an open-loop prefix before re-querying. While this interface improves local motion continuity, deployment still requires choosing the execution horizon: how much of each predicted chunk should be executed before acquiring a new observation. However, our experiments show that success is strongly task-dependent and non-monotonic with respect to the execution horizon, making a single constant horizon an unreliable deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories by identifying low-speed transition points in the predicted speed profile and using them as candidate replanning boundaries. Because PACE uses only the predicted action chunk, it is plug-and-play and requires no retraining or access to policy internals. We validate PACE through large-scale evaluations in both simulation and real-robot settings. On 50 RoboTwin2.0 tasks, PACE raises the average success rate from 57.8% to 64.2%. In real-robot experiments on bimanual ALOHA and single-arm Franka platforms, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%. Ablations and rollout-level analyses show that PACE adapts execution horizons across manipulation phases, shortening near transitions while preserving longer execution during coherent motion.
Junnan Nie, Jiayi Li, Chenghao Liu +5
Peking University · Peking University. · JD Explore Academy. +1