While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6× the success rate of PI0.5 in real-world dynamic settings and about 5× on DynBench. We will open-source all the code and weights.
Figures & tables
Figure 1: Left: Comparison on Kinetix benchmark and real-world tasks. Top Right: Current VLAs fail to perform dynamic tasks due to three gaps: the perception gap, latency gap, and control gap. Bottom Right: Our DSDyn-VLA addresses these gaps by introducing optical flow for motion perception, future state awareness for latency compensation, and a refiner for real-time correction.
Figure 2: Illustration of VLA asynchronous inference. VLA initiates a new inference when Tlat (representing the VLA inference latency steps) actions remain in the action buffer, ensuring the new action chunk is available when the buffer is exhausted. The observation for inference lags the execution by Tlat steps.
Figure 3: Architecture of our DSDyn-VLA, which consists of Flow-Planner and Res-Refiner.
Figure 4: Performance comparison on the Kinetix benchmark. Top left: Execution horizon vs. average solve rate with a fixed inference delay of 1. Bottom left: Inference delay vs. average solve rate with a fixed execution horizon = max (inference delay, 1). Right: Inference delay vs. solve rate for individual tasks. Each setting consists of 2048 trials.
Delay δ=0
Delay δ=1
Delay δ=2
Delay δ=3
Delay δ=4
Execution Horizon (K)
1
8
1
7
2
6
3
5
4
DSDyn-VLA
97.2%
97.2%
97.3%
97.1%
96.9%
97.3%
97.0%
97.3%
97.1%
w/o Future
97.0%
97.1%
93.6%
93.1%
91.6%
91.9%
88.8%
89.6%
85.6%
w/o Refiner
93.9%
90.9%
93.3%
91.7%
91.1%
91.3%
89.1%
89.1%
87.0%
Table 2: Ablation study of average success rates on Kinetix under varying latencies ( δ ) and execution horizons ( K ). We evaluate different execution sizes for each delay setting. DSDyn-VLA demonstrates superior performance compared to ablated variants.
Method
Pick
Drop
Stack
Pour
Sort
Insert
Ball
Can
GreenBall
Avg
PI0.5
12%
14%
1%
10%
19%
1%
18%
10%
2%
9.7%
RTC
16%
41%
7%
13%
25%
1%
32%
25%
15%
19.4%
VLASH
19%
55%
8%
30%
37%
12%
35%
32%
16%
27.1%
DynamicVLA
28%
55%
19%
48%
43%
17%
44%
36%
19%
34.3%
DSDyn-VLA w/o Refiner
49%
56%
32%
54%
49%
19%
46%
48%
31%
42.7%
DSDyn-VLA w/o Future
37%
58%
30%
54%
53%
24%
50%
42%
36%
42.7%
Table 3: Results on our DynBench. Each method is evaluated with 100 trials for each task.
Figure 5: Training efficiency and performance comparison on Kinetix tasks. The Res-Refiner (Ours, red) significantly outperforms the standard RL baseline (Blue), achieving 4.9x faster convergence and a +18% higher success rate .
Figure 6: Case comparison of our DSDyn-VLA and baseline (PI0.5).
Method Variant
Success Rate
DSDyn-VLA (Full)
48.5%
w/o Optical Flow
30.5%
w/o Polar Map
33.0%
w/o Future State Awareness
36.5%
w/o Res-Refiner
25.0%
Res-Refiner ( w/ SFT )
37.5%
Table 4: Design choice ablation on four real-world task (50 trials). We compare the full model against variants with alternative visual encodings or training strategies.
Table 7: Robustness under degraded optical-flow conditions, averaged over four real-world tasks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Spatial
Object
Goal
LIBERO-10
PI0.5
98.67 ± 0.23
98.30 ± 0.10
97.87 ± 0.42
92.60 ± 0.17
DSDyn-VLA
99.17 ± 0.15
98.73 ± 0.06
98.17 ± 0.29
93.30 ± 0.30
Appendix
Table 8: Success rates (%) on the four LIBERO suites (mean ± std over 3 seeds). The dynamic-oriented components of DSDyn-VLA do not degrade performance on static instruction-following tasks.
Method
SR (%)
MS
PI0.5
9.63 ± 0.08
26.19 ± 0.41
PUMA
17.18 ± 0.16
35.08 ± 0.13
DSDyn-VLA
20.74 ± 0.25
41.64 ± 0.66
Appendix
Table 9: Results on the DOMINO benchmark for dynamic manipulation (mean ± std over 3 seeds).
Type
Task
w/o Future
w/o Refiner
DSDyn-VLA
δ=0
δ=2
δ=4
δ=0
δ=2
δ=4
δ=0
δ=2
δ=4
High-Dyn
Catapult
91.2%
49.3%
41.5%
73.1%
51.7%
56.1%
90.9%
92.2%
90.6%
Walker
90.9%
75.4%
0.4%
87.4%
86.8%
86.9%
91.0%
89.9%
89.7%
Catcher_V3
99.9%
99.8%
99.7%
73.5%
70.4%
21.7%
99.7%
99.7%
99.7%
Regular
Car Launch
99.5%
99.2%
99.5%
98.9%
99.4%
99.3%
99.4%
99.5%
99.4%
Cartpole Thrust
100%
100%
100%
100%
100%
100%
100%
100%
99.9%
Appendix
Table 10: Detailed ablation results on 12 Kinetix tasks under varying latencies ( δ∈{0,2,4} ) and execution horizon of 4. We compare the full DSDyn-VLA against variants lacking future state awareness ( w/o Future ) or the residual refinement stream ( w/o Refiner ).
Setting
Avg. Success (%)
RTC
18.0
RTC + history frames + PPO
24.0
VLASH
26.5
VLASH + history frames + PPO
32.5
DynamicVLA
37.5
DynamicVLA + history frames + PPO
39.0
Appendix
Table 11: Fairer comparison with temporally strengthened baselines and ablations on motion representation. For a fair comparison, we augment RTC , VLASH , and DynamicVLA with two historical frames and 1 hour of end-to-end PPO training. We further compare several motion-input variants of our DSDyn-VLA .
Figure 7: Qualitative rollouts of DSDyn-VLA on DynBench. Top: Dynamic Stacking. Bottom: Dynamic Grasp. From left to right, the snapshots show that DSDyn-VLA tracks moving objects and completes the task under continuous motion.
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception-execution mismatch: object states continue to evolve during inference, making actions predicted from past observations stale at execution time. We present DynamicVLA, a latency-aware VLA model for dynamic object manipulation. It combines a compact 0.4B architecture and convolutional vision encoder for efficient multimodal inference with a continuous inference schedule that overlaps reasoning and execution for non-blocking control. Latent-aware Action Streaming then discards latency-invalid action prefixes and executes only the temporally valid suffix of each predicted chunk, preserving action-time alignment under dynamic object motion. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations in simulation and on real robots show that DynamicVLA improves dynamic manipulation success under changing object motion, perception-heavy instructions, and unseen motion patterns.
Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi-dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics-aware VLA architecture. By integrating scene-centric historical optical flow and specialized world queries to implicitly forecast object-centric future states, PUMA couples history-aware perception with short-horizon prediction. Results demonstrate that PUMA achieves state-of-the-art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H-EmbodVis/DOMINO.
Heng Fang, Shangru Li, Shuhan Wang +3
Huazhong University of Science and Technology, China · Huawei Technologies Co. Ltd, China
Vision-Language-Action (VLA) models generalize across static manipulation but fail when objects move during task execution. They map the current observation to an action and assume the scene is stationary between observation and execution, so at any non-trivial object speed the resulting latency exceeds the time available to grasp. We close this gap with AHEAD (Anticipatory Horizon Extrapolation with Adaptive Dynamics), a predict-then-act wrapper that augments a frozen VLA with a motion-aware latent world model. A small world model trained on manipulation video forecasts future patch tokens in the VLA's feature space, conditioned on per-token velocity and acceleration from optical flow. A language-and-motion saliency mask concentrates prediction on task-relevant patches, and the model rolls forward for an adaptive horizon, halting when prediction uncertainty crosses a threshold. The frozen action decoder then receives the predicted future tokens in place of the current ones. AHEAD adds 4.9M parameters to a frozen 7B OpenVLA and reaches 79 to 97% success across 20 dynamic simulation scenarios where the strongest baseline reaches 31 to 58%. On a physical UFactory xArm 7, AHEAD succeeds on 29/30 to 30/30 on three conveyor and rolling-ball tasks, 23/30 on paddle interception, and 19/30 on projectile catching where every baseline scores 0/30.
Shahram Najam Syed, Arthur Jakobsson, Haoran Hao +1
Robotics Institute, Carnegie Mellon University, Pittsburgh, USA