Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
Authors: Di Wu, Rongtian Shen, Ping Liu, Yan Shen, Zhenhan Yin, Shun Zuo, Xuhua Chen, He Zheng, +3 more
Organizations: Magic-Lab Team, Magiclab Robotics Inc. · Southeast University · Harbin Institute of Technology · Tongji University · Jilin University · Zhejiang University
Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-side delays mainly arise from perception acquisition, communication scheduling, and physical response. Analysis of the velocity field shows relatively stable magnitude and direction in early integration, followed by stronger directional correction near the terminal steps. Based on this stage heterogeneity, we propose two-stage non-uniform denoising, reducing the number of steps from 10 to 2 and model-inference time from 61.557 ms to 21.956 ms. We also develop a distributed real-time VLA framework with independent inference, action-publication, and robot-control rates, modular observation acquisition, and action-provenance logging. Using π0.5 as the baseline, we evaluate six real-time execution methods on a long-horizon physical garment-folding task. Legato performs best overall among training-based methods, while Temporal Smoothing leads among training-free methods; both perform strongly in task success, completion time, action continuity, and acceleration smoothness. Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance. These results motivate joint optimization of model-inference efficiency and robot-system timing.
Figures & tables
Figure 1 : Time-scale mismatch between VLA inference and robot control.
Layer
Latency component
Measured value
Interpretation
Perception
Primary-camera timestamp offset
17.7±1.8 ms
Offset of the RealSense internal timestamp relative to actual image acquisition
Perception
Primary-camera image readout
32.8±3.3 ms
Input/output (I/O) delay from the RealSense timestamp to image receipt by the host process
State
Proprioceptive feedback
26.4±6.2 ms
Delay from the physical robot state to host observation of the corresponding software development kit (SDK) feedback
Execution
Command-to-motion response
15.3±5.4 ms
Delay from command publication until physical motion follows the target phase
Model
Standard Flow, 10 NFE
61.6±6.2 ms
Latency statistics from 620 inference calls
Table 1 : Measured latency components.
Figure 2 : Joint system-latency calibration apparatus and signal-acquisition workflow. The RealSense camera records a visual time code on the display and an end-effector ArUco marker, while the host logs image timestamps, receipt times, joint commands, and proprioceptive feedback. Timestamp differences and phase relationships among periodic signals estimate the camera timestamp offset, image-readout latency, motion-response latency, and proprioceptive-feedback latency.
Figure 3 : Velocity-field trajectories under standard 10-NFE Flow sampling. Changes in velocity RMS rk (left) and the angle θk between adjacent integration steps (right) are concentrated near the endpoint.
Figure 4 : Training and inference pipeline for two-stage non-uniform action denoising. Training selects the standard Flow Matching branch or the two-stage branch with equal probability. The latter applies a large Flow update followed by short-interval mean-velocity prediction for terminal refinement. Inference reuses the same visual–language–proprioceptive condition and evaluates the action expert only twice along the non-uniform schedule 1→τ→0 . The right side compares uniform two-step sampling with the proposed asymmetric schedule.
Figure 5 : Distributed real-time VLA inference and execution framework.
Method
System layer
Principal operation
Primary objective
Naive Asynchronous
Scheduling
Decouple model inference from robot execution
Remove robot waiting during inference
Temporal Smoothing
Action buffer
Blend the old-chunk tail with the new-chunk prefix
Suppress velocity and acceleration discontinuities
Inference-time RTC
Generation constraint
Fix committed actions and complete the remaining trajectory
Reduce mismatch between old and new chunks
Training-time RTC
Policy conditioning
Simulate latency and condition on an action prefix during training
Learn low-overhead action continuation
Legato
Learned generation
Learn native action-continuation dynamics
Internalize inter-chunk continuity
VLASH
State alignment
Propagate state to the expected execution time
Reduce prediction error caused by stale state
Table 2 : Layered taxonomy of six real-time VLA execution strategies.
Figure 8
Figure 8: Stage sequence of the bimanual garment-folding task.
Method
Inference (Hz)
Publication (Hz)
NFE
Success rate
95% Wilson interval
Mean time (s)
Throughput (h -1 )
VLASH
2
30
10
28/30 (93.3%)
[78.7, 98.2]
77.73
43.22
Training-time RTC
1
30
10
19/30 (63.3%)
[45.5, 78.1]
129.37
17.62
Legato
2
50
5
29/30 (96.7%)
[83.3, 99.4]
73.56
47.31
Temporal Smoothing
3
30
10
23/30 (76.7%)
[59.1, 88.2]
93.93
29.38
Naive Asynchronous
1
30
10
19/30 (63.3%)
[45.5, 78.1]
124.13
18.37
Inference-time RTC
3
30
10
18/30 (60.0%)
[42.3, 75.4]
133.73
16.15
Table 3 : Task-level results for real-time VLA execution methods.
Method
Maximum acceleration (rad/s 2 )
Tail Gap (rad)
Switch Gap (rad)
Naive Asynchronous
744.1
0.1799
0.1538
Temporal Smoothing
25.6
0.0967
0.0998
Inference-time RTC
65.55
0.1310
0.1164
Training-time RTC
392.8
0.1201
0.0630
Legato
46.6
0.0193
0.0062
VLASH
73.9
0.0459
0.0173
Table 4 : Action-continuity results for left_j3 ; lower is better for every metric.
Figure 9 : Local left_j3 trajectory for Naive Asynchronous Execution.
Figure 10 : Local left_j3 trajectory for Legato.
Sampler
NFE
Mean (ms)
Speedup
Time reduction
Standard Flow
10
61.557
1.000×
0
Two-stage non-uniform
2
21.956
2.804×
64.33%
Table 5 : Inference time for standard Flow and two-stage denoising.
Data
Model
Joint MAE (rad)
Joint RMSE (rad)
Left TCP (mm)
Right TCP (mm)
Offline sequence I
Standard Flow
0.008525
0.012877
11.396
7.676
Two-stage
0.008396
0.013329
11.493
6.461
Offline sequence II
Standard Flow
0.006503
0.008931
6.930
5.684
Two-stage
0.006797
0.009104
4.831
6.718
Table 6 : Action-error comparison.
Model efficiency
Task performance
Method
Config.
NFE
Inference time (ms) ↓
Relative speedup ↑
Success rate (%) ↑
Successful-trial time (s) ↓
Legato
Baseline
5
36.808
1.000×
29/30 (96.7)
73.56
Legato + two-stage 2-NFE
Combined
2
21.956
1.676×
26/30 (86.7)
76.13
Temporal Smoothing
Baseline
10
61.557
1.000×
23/30 (76.7)
93.93
Temporal Smoothing + two-stage 2-NFE
Combined
2
21.956
2.804×
21/30 (70.0)
98.63
Table 7 : Combined evaluation of two-stage 2-NFE denoising with Legato and Temporal Smoothing. ↑ and ↓ indicate that higher and lower values are better, respectively.
Figure 11 : Chronological development and methodological taxonomy of representative real-time VLA studies. The six branches denote asynchronous inference and execution-time alignment, inter-chunk continuity and action priors, adaptive execution horizons and multi-rate control, fast action generation and streaming inference, continuous-time action representations and physical executability, and execution-time correction and system-level deployment. Only representative studies are shown; publication dates and version information follow the cited references.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Garment condition
Successes
Mean time (s)
Throughput (h -1 )
VLASH
Small purple
10/10
56.4
63.83
Large pink
9/10
99.2
32.66
Medium red
9/10
77.6
41.75
Training-time RTC
Small purple
8/10
116.5
24.72
Large pink
7/10
124.4
20.26
Medium red
4/10
147.2
9.78
Appendix
Table S1: Task-level results across garment conditions. Each row contains ten physical trials, and throughput is the number of successful tasks completed per hour.
Hyperparameter
Value
Training dataset
3,048 teleoperated demonstration trajectories
Training steps
30,000
Batch size
16
Optimizer
AdamW, β1=0.9 , β2=0.95 , ϵ=10−8
Learning-rate schedule
Peak learning rate 2.5×10−5 ; linear warmup for 1,000 steps, followed by cosine decay to 2.5×10−6
Regularization and gradient clipping
Weight decay 0.01 ; maximum gradient norm 1.0
Appendix
Table S2: Training configuration for two-stage non-uniform denoising.
Figure S1: Component-wise RMS comparison between predicted and Flow Matching GT velocity vectors in a representative validation sequence. Colored solid curves show predicted velocity RMS; black horizontal dashed lines show the corresponding GT velocity RMS.
Table S3: Supplementary action-error metrics on the fixed validation sequence.
Figure S2: Per-joint action errors on the fixed validation sequence. (A) MAE for 12 rotational joints. (B) Left and right gripper-command MAE. (C) Frame-wise mean absolute error over the 12 rotational joints. Blue denotes the standard 10-NFE model, and orange denotes the two-stage 2-NFE model.
Figure S3: Per-joint comparison.
Figure S4: Bimanual TCP trajectories and pose errors on the fixed validation sequence. (A,D) Three-dimensional left and right TCP trajectories. (B,E) Euclidean translation error. (C,F) SO(3) geodesic orientation error. Black denotes the TCP trajectory obtained by applying URDF forward kinematics to the recorded action. Blue and orange denote the standard 10-NFE and two-stage 2-NFE models.
Figure S5: Reference inference and execution modes.