Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
Organizations: Magic-Lab Team, Magiclab Robotics Inc. · Southeast University · Harbin Institute of Technology · Tongji University · Jilin University · Zhejiang University
Abstract
Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-side delays mainly arise from perception acquisition, communication scheduling, and physical response. Analysis of the velocity field shows relatively stable magnitude and direction in early integration, followed by stronger directional correction near the terminal steps. Based on this stage heterogeneity, we propose two-stage non-uniform denoising, reducing the number of steps from 10 to 2 and model-inference time from 61.557 ms to 21.956 ms. We also develop a distributed real-time VLA framework with independent inference, action-publication, and robot-control rates, modular observation acquisition, and action-provenance logging. Using π0.5 as the baseline, we evaluate six real-time execution methods on a long-horizon physical garment-folding task. Legato performs best overall among training-based methods, while Temporal Smoothing leads among training-free methods; both perform strongly in task success, completion time, action continuity, and acceleration smoothness. Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance. These results motivate joint optimization of model-inference efficiency and robot-system timing.
Figures & tables
| Layer | Latency component | Measured value | Interpretation |
| Perception | Primary-camera timestamp offset | ms | Offset of the RealSense internal timestamp relative to actual image acquisition |
| Perception | Primary-camera image readout | ms | Input/output (I/O) delay from the RealSense timestamp to image receipt by the host process |
| State | Proprioceptive feedback | ms | Delay from the physical robot state to host observation of the corresponding software development kit (SDK) feedback |
| Execution | Command-to-motion response | ms | Delay from command publication until physical motion follows the target phase |
| Model | Standard Flow, 10 NFE | ms | Latency statistics from inference calls |
| Method | System layer | Principal operation | Primary objective |
| Naive Asynchronous | Scheduling | Decouple model inference from robot execution | Remove robot waiting during inference |
| Temporal Smoothing | Action buffer | Blend the old-chunk tail with the new-chunk prefix | Suppress velocity and acceleration discontinuities |
| Inference-time RTC | Generation constraint | Fix committed actions and complete the remaining trajectory | Reduce mismatch between old and new chunks |
| Training-time RTC | Policy conditioning | Simulate latency and condition on an action prefix during training | Learn low-overhead action continuation |
| Legato | Learned generation | Learn native action-continuation dynamics | Internalize inter-chunk continuity |
| VLASH | State alignment | Propagate state to the expected execution time | Reduce prediction error caused by stale state |
| Method | Inference (Hz) | Publication (Hz) | NFE | Success rate | 95% Wilson interval | Mean time (s) | Throughput (h -1 ) |
| VLASH | 2 | 30 | 10 | 28/30 (93.3%) | [78.7, 98.2] | 77.73 | 43.22 |
| Training-time RTC | 1 | 30 | 10 | 19/30 (63.3%) | [45.5, 78.1] | 129.37 | 17.62 |
| Legato | 2 | 50 | 5 | 29/30 (96.7%) | [83.3, 99.4] | 73.56 | 47.31 |
| Temporal Smoothing | 3 | 30 | 10 | 23/30 (76.7%) | [59.1, 88.2] | 93.93 | 29.38 |
| Naive Asynchronous | 1 | 30 | 10 | 19/30 (63.3%) | [45.5, 78.1] | 124.13 | 18.37 |
| Inference-time RTC | 3 | 30 | 10 | 18/30 (60.0%) | [42.3, 75.4] | 133.73 | 16.15 |
| Method | Maximum acceleration (rad/s 2 ) | Tail Gap (rad) | Switch Gap (rad) |
| Naive Asynchronous | 744.1 | 0.1799 | 0.1538 |
| Temporal Smoothing | 25.6 | 0.0967 | 0.0998 |
| Inference-time RTC | 65.55 | 0.1310 | 0.1164 |
| Training-time RTC | 392.8 | 0.1201 | 0.0630 |
| Legato | 46.6 | 0.0193 | 0.0062 |
| VLASH | 73.9 | 0.0459 | 0.0173 |
| Sampler | NFE | Mean (ms) | Speedup | Time reduction |
| Standard Flow | 10 | 61.557 | 0 | |
| Two-stage non-uniform | 2 | 21.956 | 64.33% |
| Data | Model | Joint MAE (rad) | Joint RMSE (rad) | Left TCP (mm) | Right TCP (mm) |
| Offline sequence I | Standard Flow | 0.008525 | 0.012877 | 11.396 | 7.676 |
| Two-stage | 0.008396 | 0.013329 | 11.493 | 6.461 | |
| Offline sequence II | Standard Flow | 0.006503 | 0.008931 | 6.930 | 5.684 |
| Two-stage | 0.006797 | 0.009104 | 4.831 | 6.718 |
| Model efficiency | Task performance | |||||
| Method | Config. | NFE | Inference time (ms) | Relative speedup | Success rate (%) | Successful-trial time (s) |
| Legato | Baseline | 5 | 29/30 (96.7) | 73.56 | ||
| Legato + two-stage 2-NFE | Combined | 2 | 26/30 (86.7) | 76.13 | ||
| Temporal Smoothing | Baseline | 10 | 23/30 (76.7) | 93.93 | ||
| Temporal Smoothing + two-stage 2-NFE | Combined | 2 | 21/30 (70.0) | 98.63 | ||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Garment condition | Successes | Mean time (s) | Throughput (h -1 ) |
| VLASH | Small purple | 10/10 | 56.4 | 63.83 |
| Large pink | 9/10 | 99.2 | 32.66 | |
| Medium red | 9/10 | 77.6 | 41.75 | |
| Training-time RTC | Small purple | 8/10 | 116.5 | 24.72 |
| Large pink | 7/10 | 124.4 | 20.26 | |
| Medium red | 4/10 | 147.2 | 9.78 |
| Hyperparameter | Value |
| Training dataset | teleoperated demonstration trajectories |
| Training steps | |
| Batch size | |
| Optimizer | AdamW, , , |
| Learning-rate schedule | Peak learning rate ; linear warmup for steps, followed by cosine decay to |
| Regularization and gradient clipping | Weight decay ; maximum gradient norm |
| Metric | Standard , 10 NFE | Two-stage, 2 NFE |
| 12-joint 95th-percentile absolute error (P95, rad) | 0.019209 | 0.018704 |
| 12-joint maximum absolute error (rad) | 0.057459 | 0.052449 |
| Left/right gripper MAE (command unit) | 0.000482 / 0.000634 | 0.000548 / 0.000344 |
| Left/right TCP rotation MAE ( ∘ ) | 1.047 / 0.902 | 0.911 / 1.169 |