Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies. We exploit this asymmetry to introduce Urgency-Aware Denoising (UAD), a novel inference-time framework that allocates denoising computation according to when each action is physically needed. UAD releases time-critical urgent actions after fewer denoising steps while overlapping the continued background refinement of tail actions with physical execution. However, heterogeneous denoising introduces two key challenges: early-release errors in urgent actions and trajectory inconsistency in tail actions. UAD elegantly resolves both through two core mechanisms: Trajectory Reconciliation, which reconstructs unified internal state evolution to restore joint denoising coherence without additional model evaluations, and Ghost Action Correction, which leverages non-executed ghost continuations to dynamically compensate for early-release errors across remaining executable actions. Extensive evaluations across multiple VLA architectures, simulation benchmarks, and real-world manipulation tasks demonstrate that UAD achieves up to a 1.89x speedup in average action availability latency while maintaining comparable success rates to vanilla inference with optimal denoising budget, offering a more favorable success-latency trade-off than state-of-the-art VLA acceleration baselines.
Figures & tables
Figure 1: Illustration of Urgency-Aware Denoising (UAD). Insight: Actions are generated jointly but consumed sequentially, creating heterogeneous execution urgencies. Method: UAD releases urgent actions after fewer denoising steps and overlaps the remaining denoising with physical execution. Modules: Urgency-Aware Early Release (UAER) releases urgent actions early to reduce latency, Trajectory Reconciliation (TR) restores consistent joint denoising, and Ghost Action Correction (GAC) compensates for early-release errors using ghost actions. Benefits: UAD reduces action availability latency while maintaining comparable task success rates.
Figure 2: Overview of UAD
Figure 3: Relative L2 error to vanilla inference ( N∗=3 ). TR reduces tail action error while ghost states remain close to vanilla.
SmolVLA + LIBERO-Plus
SmolVLA + Meta-World+
π0.5 + Meta-World+
Method
SR
AAL
Δ SR
Speedup
SR
AAL
Δ SR
Speedup
SR
AAL
Δ SR
Speedup
(%) ↑
(ms) ↓
(pp) ↑
( × ) ↑
(%) ↑
(ms) ↓
(pp) ↑
( × ) ↑
(%) ↑
(ms) ↓
(pp) ↑
( × ) ↑
Vanilla ( N∗ )
35.3
21.6
0.0
1.00
70.8
14.4
0.0
1.00
70.5
11.2
0.0
1.00
Vanilla ( N=1 )
27.3
11.2
−8.0
1.93
60.1
10.8
−10.7
1.33
66.6
9.0
−3.9
1.24
D3P
36.3
17.6
+1.0
1.23
70.3
13.9
−0.5
1.04
70.5
11.6
0.0
0.97
RTC
8.8
2.7
−26.5
8.00
52.3
1.3
−18.5
11.08
59.8
1.5
−10.7
7.47
Table 1: Main results across three simulation settings. See Metrics for the definitions of metrics.
Figure 4: Success rate-latency trade-offs across three simulation settings. UAD achieves low AAL while retaining high SR. Broken horizontal axes accommodate RTC’s substantially lower AAL. Error bars indicate Wilson 95% CIs for SR, and episode bootstrap 95% CIs for AAL. Some AAL CIs are too narrow to be visible at the plotted scale.
Figure 6Figure 7
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Task success rate (red circles, left axis) and action availability latency (blue triangles, right axis) for different numbers of denoising steps N across six model-benchmark settings.
Figure 9: Average tail blocking time for different values of U at K=1 .
Figure 10: Real-world evaluation tasks and camera observations. We evaluate five manipulation tasks using front lateral and wrist-mounted camera views. Camera views are cropped.
SmolVLA
SmolVLA
π0.5
Real-World
LIBERO-Plus
Meta-World+
Meta-World+
Deployment
Method
AAL
DBL
AAL
DBL
AAL
DBL
AAL
DBL
(ms) ↓
(ms) ↓
(ms) ↓
(ms) ↓
(ms) ↓
(ms) ↓
(ms) ↓
(ms) ↓
Vanilla ( N∗ )
21.6
120.3
14.4
52.4
11.2
47.1
34.3
216.6
Vanilla ( N=1 )
11.2
18.2
10.8
17.9
9.0
24.6
14.9
19.9
D3P
17.6
122.2
13.9
53.7
11.6
36.1
-
-
Appendix
Table 3: Action availability latency (AAL) and DiT blocking latency (DBL) for the evaluated settings and methods.
SmolVLA
π0.5
Method
Easy
Medium
Hard
Very hard
Overall
Easy
Medium
Hard
Very hard
Overall
Vanilla ( N∗ )
85.2
56.4
50.0
47.0
70.8
80.7
58.6
68.3
42.0
70.5
Vanilla ( N=1 )
76.1
45.5
38.3
29.0
60.1
79.8
54.1
55.0
34.0
66.6
D3P
84.1
57.7
50.8
44.0
70.3
80.2
61.8
68.3
38.0
70.5
RTC
63.2
47.7
30.8
27.0
52.3
69.8
52.3
49.2
33.0
59.8
BAC
80.4
45.9
38.3
31.0
62.8
75.9
39.5
44.2
20.0
58.5
Appendix
Table 4: Success rate (%) on Meta-World+ by task difficulty (easy / medium / hard / very hard: 28 / 11 / 6 / 5 tasks, 20 episodes each).
Method
Button
Grapes
Flower
Tidy
Stack
Overall
Vanilla ( N=10 )
100
85
65
95
75
84
Vanilla ( N=1 )
100
70
55
75
55
71
UAD (ours)
100
90
65
90
75
84
Appendix
Table 5: Success rate (%) on individual real-world manipulation tasks. Each task is evaluated over 20 episodes. Detailed task descriptions are provided in Fig. 10 .
Figure 11: Success rate-latency trade-offs across three additional simulation settings. Broken horizontal axes accommodate RTC’s substantially lower AAL. Error bars indicate Wilson 95% CIs for SR, and episode bootstrap 95% CIs for AAL. Some AAL CIs are too narrow to be visible at the plotted scale.
Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-side delays mainly arise from perception acquisition, communication scheduling, and physical response. Analysis of the velocity field shows relatively stable magnitude and direction in early integration, followed by stronger directional correction near the terminal steps. Based on this stage heterogeneity, we propose two-stage non-uniform denoising, reducing the number of steps from 10 to 2 and model-inference time from 61.557 ms to 21.956 ms. We also develop a distributed real-time VLA framework with independent inference, action-publication, and robot-control rates, modular observation acquisition, and action-provenance logging. Using π0.5 as the baseline, we evaluate six real-time execution methods on a long-horizon physical garment-folding task. Legato performs best overall among training-based methods, while Temporal Smoothing leads among training-free methods; both perform strongly in task success, completion time, action continuity, and acceleration smoothness. Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance. These results motivate joint optimization of model-inference efficiency and robot-system timing.
Di Wu, Rongtian Shen, Ping Liu +8
Magic-Lab Team, Magiclab Robotics Inc. · Southeast University · Harbin Institute of Technology +3
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, thereby drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes or even different tasks. Experimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving inference acceleration of up to 11.75× and 34.43× for representative flow-based VLA models, π0.5 and GR00T-N1.6, respectively.