Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy's reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA's action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA's action distribution. Our resulting method matches the original VLA's success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.
Figures & tables
Figure 1: Asynchronous execution and distribution alignment. The point clouds schematically represent different action distributions at t , with each point corresponding to a possible action sample; inference of such a sample takes Δt . (A) Synchronous execution queries the VLA at time t , so execution has to wait until the action is generated at t+Δt . (B–C) Asynchronous execution instead starts generation at t−Δt , while the robot carries out its committed actions, so that the next chunk is prepared for execution at t . (B) Real-Time Chunking produces a narrower subset of the actions compared to the VLA when the training data is non-Markovian. (C) Recursive Flow-Field Distillation trains the asynchronous policy using the VLA’s action-generation flow, recovering a broader proposal distribution. (D) Propose–Resolve uses the latest observation to estimate each candidate’s likelihood under the corresponding VLA distribution and selects the highest-scoring continuation.
Figure 2: Execution timing and action distributions at handoff. Schematic with d=3 and s=5 . Left: Synchronous execution queries the VLA at time t for the five actions executed over [t,t+s] . Asynchronous execution instead begins generation at t−d : the first d actions form the committed prefix P , while the following s actions form the continuation U executed after handoff. Note that the schematic assumes prefix-compatible generation, so U begins from the state reached after executing P ; enforcing this compatibility is the original objective from RTC. Right: The resulting distributions over the same s -step execution interval. Black curves show possible VLA action chunks conditioned on o+ , while red curves show asynchronous continuations generated from (o−,P) .
Figure 3: One ToolHang handoff at d=10 .
Figure 4: Distribution comparison at o+ . For each handoff tuple (o−,P,o+) at d=10 , we draw 128 continuations from each asynchronous method conditioned on (o−,P) and independent banks of 128 support and 128 query continuations from πVLA(⋅∣o+) . Left: Per-state precision and recall over N=2,775 ToolHang demonstration handoffs ( 291 for RTC), shown as histograms. Right: Mean precision and recall across RoboMimic and LIBERO, evaluated on both demonstration states and states from VLA rollouts. Appendix D details how precision and recall are computed, extends this analysis across delays d , and additionally reports the distribution obtained under the teacher-sample supervision introduced in Section 4.1 .
Figure 5: Real-time execution success and ablation. Top: Rollout success on LIBERO and RoboMimic across asynchronous delays, compared with the synchronous VLA reference. Our method uses K=32 RFD proposals followed by Propose–Resolve. Bottom: Ablation at d=10 , comparing each asynchronous proposer with a single continuation ( K=1 ) and with Propose–Resolve over K=32 candidates.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Distribution alignment across delays. The support comparison from Figure 4 , extended across delays d=0,…,10 on ToolHang, Square, and Transport, showing how relative recall and precision evolve on demonstration and rollout states.
Figure 7: Action support with teacher-sample supervision. The ToolHang comparison from Figure 4 , now including teacher-sample supervision (TSS) and reporting precision and recall relative to each state’s VLA self-coverage.
Hyperparameter
Base VLA / TT-RTC / RFD
Resolver
Optimizer
AdamW
AdamW
Adam (β1,β2)
(0.9,0.95)
(0.9,0.95)
Weight decay
10−10
10−4
Gradient clipping norm
10
5
Learning-rate schedule
Cosine with warm-up
Cosine with warm-up
Warm-up steps
1,000
500
Appendix
Table 1: Optimization settings. Shared settings for base-VLA fine-tuning, TT-RTC, RFD, and resolver training.
Component
Task
Peak LR
Batch
Updates
Seed
Base VLA
Square
10−4
16
60,000
1000
Base VLA
Transport
10−4
8
60,000
1000
Base VLA
ToolHang
10−4
16
60,000
1000
Base VLA
LIBERO
—
—
Released checkpoint
—
TT-RTC
Square
10−4
16
100,000
1000
TT-RTC
Transport
2.5×10−5
16
25,000
1000
Appendix
Table 2: Training configurations. Policy update counts extend through the evaluated checkpoint; resolver counts give the full training budget. Learning rates are configured peaks. Arrows and sums denote successive stages.
Model
K
Propose mean / p99
Resolve mean / p99
Total mean
SmolVLA
1
27.0 / 27.6
0 / 0
27.0
SmolVLA
8
39.8 / 40.4
12.0 / 12.4
51.8
SmolVLA
16
56.5 / 57.1
12.0 / 12.5
68.6
SmolVLA
32
90.8 / 91.4
12.0 / 12.3
102.8
Appendix
Table 3: Compute latency of Propose–Resolve. RFD proposals with energy-argmax resolution on an NVIDIA GeForce RTX 4090, using 50-action chunks and ten flow steps. All times are in milliseconds. Proposal generation overlaps with execution; resolution follows observation of o+ . At K=1 , resolution is bypassed. Total means sum separately measured proposal and resolution means; they are not joint pipeline measurements.
Vision-Language-Action (VLA) policies increasingly rely on asynchronous inference to hide large-model latency behind ongoing robot motion. While this avoids the stop-and-go behavior of synchronous action-chunk execution, it creates a prediction-execution mismatch: the next chunk is computed from a stale observation at inference start but executed only after the robot and scene have evolved. As a result, actions that fit the prediction-time state can become misaligned with the execution-time state. Existing runtime repair, behavior-cloning, and preference-alignment approaches do not directly teach the policy to resolve this stale-input mismatch. We propose DEFLECT, an offline post-training framework for delay-robust asynchronous VLAs. DEFLECT converts latency-induced mismatch into counterfactual preference supervision: a frozen reference VLA generates a preferred chunk from the future execution-time observation and a rejected chunk from the stale prediction-time observation. The trainable policy scores both chunks under the same deployment-time input, learning to favor execution-time-aligned actions while a supervised fine-tuning anchor preserves the expert action manifold. DEFLECT requires no human preference labels, reward models, online robot rollouts, architectural changes, or additional inference-time computation. Across Kinetix, LIBERO, and three real-robot tasks, DEFLECT improves delay robustness over strong asynchronous VLA baselines, raising high-latency success by up to 6.4 percentage points and achieving a 4.6 percentage-point gain at the longest delay on a real-scale VLA.
Yixiang Zhu, Yonghao Chen, Zijie Yang +2
The Hong Kong University of Science and Technology (Guangzhou) · One Robotics
Real-time deployment of Vision-Language-Action (VLA) policies necessitates asynchronous execution, wherein subsequent action chunks are computed concurrently with the execution of the current chunk, leading to prediction-execution misalignment and manifesting as inter-chunk discontinuities. Existing methods either superficially smooth chunk boundaries, require costly policy optimization, or exclusively forward-predict proprioceptive states yet neglect critical visual observations. In this paper, we propose \textbf{FutureRTC}, a plug-and-play adaptation framework that predicts execution-time observations and states for asynchronous VLA control without modifying the underlying policy. Specifically, FutureRTC features a state correction module to compensate for the discrepancy between rolled-forward and actual execution-time proprioceptive states and an observation prediction module that forecasts execution-time visual representations by leveraging robot motion as an explicit physical prior through motion-aware feature transport and reconstruction. Furthermore, we introduce a policy consistency loss to align the action chunks generated from predicted contexts with those produced under the expected execution-time inputs of the VLA policy. Extensive experiments across simulated and real-world environments demonstrate that FutureRTC achieves superior robustness to inference delays, resulting in smoother trajectories, faster execution, and consistently higher task success rates.
Hai Jiang, Yixian Zou, Binbin Liang +3
School of Aeronautics and Astronautics, Sichuan University · University of Electronic Science and Technology of China · University of Alberta
Deploying robot policies in the physical world requires satisfying two fundamental desiderata: reliability and smooth real-time execution. However, deploying state-of-the-art generalist models presents challenges on both fronts. Achieving the precision and robustness required for real-world deployment necessitates sample-efficient online reinforcement learning (RL) to adapt pretrained models. Meanwhile, the increasing scale of robot foundation models has led to higher inference latency. To satisfy real-time constraints under high latency, modern systems adopt asynchronous inference with action chunking, overlapping policy computation with chunk execution to hide latency and enable smooth control. Despite their complementary roles, integrating asynchronous execution with gradient-based online RL remains underexplored. We present SmoothRL, an online RL framework that fine-tunes a pretrained policy within an asynchronous inference loop. SmoothRL follows a value-gradient paradigm, directly updating policy parameters using gradients of the action-value function with respect to policy actions. To enable correct optimization under asynchronous execution, SmoothRL explicitly models the asynchronous inference process during training. Specifically, each generated action chunk is partitioned by frame index into three regions: a committed region, consisting of actions committed by the previous inference cycle; an execution region, containing newly generated actions executed by the robot; and a discarded region, containing actions superseded by the next inference cycle. Gradients are propagated only through the execution region, ensuring policy optimization aligns with the trajectory distribution induced by asynchronous execution. We evaluate SmoothRL on real-world robotic tasks requiring high precision, as well as highly dynamic tasks that necessitate asynchronous execution.