While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6× the success rate of PI0.5 in real-world dynamic settings and about 5× on DynBench. We will open-source all the code and weights.
Figures & tables
Figure 1: Left: Comparison on Kinetix benchmark and real-world tasks. Top Right: Current VLAs fail to perform dynamic tasks due to three gaps: the perception gap, latency gap, and control gap. Bottom Right: Our DSDyn-VLA addresses these gaps by introducing optical flow for motion perception, future state awareness for latency compensation, and a refiner for real-time correction.
Figure 2: Illustration of VLA asynchronous inference. VLA initiates a new inference when Tlat (representing the VLA inference latency steps) actions remain in the action buffer, ensuring the new action chunk is available when the buffer is exhausted. The observation for inference lags the execution by Tlat steps.
Figure 3: Architecture of our DSDyn-VLA, which consists of Flow-Planner and Res-Refiner.
Figure 4: Performance comparison on the Kinetix benchmark. Top left: Execution horizon vs. average solve rate with a fixed inference delay of 1. Bottom left: Inference delay vs. average solve rate with a fixed execution horizon = max (inference delay, 1). Right: Inference delay vs. solve rate for individual tasks. Each setting consists of 2048 trials.
Delay δ=0
Delay δ=1
Delay δ=2
Delay δ=3
Delay δ=4
Execution Horizon (K)
1
8
1
7
2
6
3
5
4
DSDyn-VLA
97.2%
97.2%
97.3%
97.1%
96.9%
97.3%
97.0%
97.3%
97.1%
w/o Future
97.0%
97.1%
93.6%
93.1%
91.6%
91.9%
88.8%
89.6%
85.6%
w/o Refiner
93.9%
90.9%
93.3%
91.7%
91.1%
91.3%
89.1%
89.1%
87.0%
Table 2: Ablation study of average success rates on Kinetix under varying latencies ( δ ) and execution horizons ( K ). We evaluate different execution sizes for each delay setting. DSDyn-VLA demonstrates superior performance compared to ablated variants.
Method
Pick
Drop
Stack
Pour
Sort
Insert
Ball
Can
GreenBall
Avg
PI0.5
12%
14%
1%
10%
19%
1%
18%
10%
2%
9.7%
RTC
16%
41%
7%
13%
25%
1%
32%
25%
15%
19.4%
VLASH
19%
55%
8%
30%
37%
12%
35%
32%
16%
27.1%
DynamicVLA
28%
55%
19%
48%
43%
17%
44%
36%
19%
34.3%
DSDyn-VLA w/o Refiner
49%
56%
32%
54%
49%
19%
46%
48%
31%
42.7%
DSDyn-VLA w/o Future
37%
58%
30%
54%
53%
24%
50%
42%
36%
42.7%
Table 3: Results on our DynBench. Each method is evaluated with 100 trials for each task.
Figure 5: Training efficiency and performance comparison on Kinetix tasks. The Res-Refiner (Ours, red) significantly outperforms the standard RL baseline (Blue), achieving 4.9x faster convergence and a +18% higher success rate .
Figure 6: Case comparison of our DSDyn-VLA and baseline (PI0.5).
Method Variant
Success Rate
DSDyn-VLA (Full)
48.5%
w/o Optical Flow
30.5%
w/o Polar Map
33.0%
w/o Future State Awareness
36.5%
w/o Res-Refiner
25.0%
Res-Refiner ( w/ SFT )
37.5%
Table 4: Design choice ablation on four real-world task (50 trials). We compare the full model against variants with alternative visual encodings or training strategies.
Table 7: Robustness under degraded optical-flow conditions, averaged over four real-world tasks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Spatial
Object
Goal
LIBERO-10
PI0.5
98.67 ± 0.23
98.30 ± 0.10
97.87 ± 0.42
92.60 ± 0.17
DSDyn-VLA
99.17 ± 0.15
98.73 ± 0.06
98.17 ± 0.29
93.30 ± 0.30
Appendix
Table 8: Success rates (%) on the four LIBERO suites (mean ± std over 3 seeds). The dynamic-oriented components of DSDyn-VLA do not degrade performance on static instruction-following tasks.
Method
SR (%)
MS
PI0.5
9.63 ± 0.08
26.19 ± 0.41
PUMA
17.18 ± 0.16
35.08 ± 0.13
DSDyn-VLA
20.74 ± 0.25
41.64 ± 0.66
Appendix
Table 9: Results on the DOMINO benchmark for dynamic manipulation (mean ± std over 3 seeds).
Type
Task
w/o Future
w/o Refiner
DSDyn-VLA
δ=0
δ=2
δ=4
δ=0
δ=2
δ=4
δ=0
δ=2
δ=4
High-Dyn
Catapult
91.2%
49.3%
41.5%
73.1%
51.7%
56.1%
90.9%
92.2%
90.6%
Walker
90.9%
75.4%
0.4%
87.4%
86.8%
86.9%
91.0%
89.9%
89.7%
Catcher_V3
99.9%
99.8%
99.7%
73.5%
70.4%
21.7%
99.7%
99.7%
99.7%
Regular
Car Launch
99.5%
99.2%
99.5%
98.9%
99.4%
99.3%
99.4%
99.5%
99.4%
Cartpole Thrust
100%
100%
100%
100%
100%
100%
100%
100%
99.9%
Appendix
Table 10: Detailed ablation results on 12 Kinetix tasks under varying latencies ( δ∈{0,2,4} ) and execution horizon of 4. We compare the full DSDyn-VLA against variants lacking future state awareness ( w/o Future ) or the residual refinement stream ( w/o Refiner ).
Setting
Avg. Success (%)
RTC
18.0
RTC + history frames + PPO
24.0
VLASH
26.5
VLASH + history frames + PPO
32.5
DynamicVLA
37.5
DynamicVLA + history frames + PPO
39.0
Appendix
Table 11: Fairer comparison with temporally strengthened baselines and ablations on motion representation. For a fair comparison, we augment RTC , VLASH , and DynamicVLA with two historical frames and 1 hour of end-to-end PPO training. We further compare several motion-input variants of our DSDyn-VLA .
Figure 7: Qualitative rollouts of DSDyn-VLA on DynBench. Top: Dynamic Stacking. Bottom: Dynamic Grasp. From left to right, the snapshots show that DSDyn-VLA tracks moving objects and completes the task under continuous motion.