Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception-execution mismatch: object states continue to evolve during inference, making actions predicted from past observations stale at execution time. We present DynamicVLA, a latency-aware VLA model for dynamic object manipulation. It combines a compact 0.4B architecture and convolutional vision encoder for efficient multimodal inference with a continuous inference schedule that overlaps reasoning and execution for non-blocking control. Latent-aware Action Streaming then discards latency-invalid action prefixes and executes only the temporally valid suffix of each predicted chunk, preserving action-time alignment under dynamic object motion. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations in simulation and on real robots show that DynamicVLA improves dynamic manipulation success under changing object motion, perception-heavy instructions, and unseen motion patterns.
Figures & tables
Figure 2: The DynamicVLA architecture. Multimodal features condition a diffusion-based action expert for action-chunk prediction.
Figure 3: Automatic Simulation and Real-world Data Collection. Environment Setup: simulation and real-world settings share objects, scenes, and synchronized cameras. Object State Acquisition: simulation provides ground-truth 6D object states, while real-world RGB observations are converted into a real-world “simulator” interface that enables automatic dynamic-manipulation data collection without teleoperation.
Methods
Interaction
Perception
Generalization
Average
CR
DA
LS
VU
SR
MP
VG
MG
DR
Success ↑
Time ↓
Diffusion Policy [ 29 ]
0.50
0.50
0.00
1.00
0.00
0.00
1.00
0.50
0.00
0.38
10.89
ACT [ 16 ]
2.50
1.50
0.00
1.00
1.50
0.50
2.50
2.50
0.00
1.33
10.80
OpenVLA-OFT [ 54 ]
3.50
0.50
0.50
0.00
1.50
0.50
3.50
2.00
0.00
1.33
10.83
π0.5 [ 55 ]
9.50
17.50
3.50
5.00
12.50
9.00
5.00
19.50
18.00
11.06
10.62
RTC [ 17 ]
12.00
15.00
4.00
7.50
13.50
8.50
6.50
14.00
20.00
11.22
10.63
Table 1: Dynamic Object Manipulation Simulation Benchmark Results. Average success rates (%) are reported across nine evaluation sub-dimensions, organized under three categories: Interaction, Perception, and Generalization. In addition, overall average success rate (Success, %) and task completion time (Time, seconds) are reported. Each method is evaluated over 1,800 trials (10 scenes × 9 dimensions × 20 trials). All baseline models are fine-tuned on the DOM dataset using their official implementations and released pretrained weights. Best results are highlighted in bold.
Figure 4: Real-world Evaluation. We compare VLA models on 16 real-world dynamic manipulation tasks using Franka and PiPER as the manipulation robot. Each method–task pair is evaluated over 60 trials, comprising 20 trials for each of three paired motion–position configurations. An additional robot arm initializes object motion, providing comparable motion conditions across trials.
Table 5
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
t−3
t−2
t−1
t
SR ↑
T.Time ↓
I.Time ↓
✗
✗
✗
✓
38.22
9.52
0.225
✗
✗
✓
✓
43.39
8.77
0.226
✗
✓
✗
✓
47.06
8.53
0.226
✓
✗
✗
✓
46.89
8.51
0.226
✗
✓
✓
✓
47.11
8.46
0.228
✓
✓
✓
✓
47.06
8.53
0.229
Appendix
Table 4: Ablation on Temporal Visual Context. The temporal observation window is varied by enabling different visual frames at time steps {t−3,t−2,t−1,t} , while keeping the model architecture, inference frequency, and execution pipeline fixed. Note that SR, T.Time, and I.Time represent the success rate (in %), task completion time (in seconds), and inference time (in seconds, measured on an NVIDIA RTX A6000 GPU), respectively.
#Layers
SR ↑
T.Time ↓
I.Time ↓
#Param ↓
8
44.17
8.92
0.127
303
16
47.06
8.53
0.226
430
24
48.44
8.43
0.317
558
32
42.11
8.39
0.373
685
Appendix
Table 5: Ablation on LLM Depth. Different LLM depths are evaluated by retaining the first l transformer layers. Note that SR, T.Time, I.Time, and #Param denote success rate (%), task completion time (seconds), inference time (seconds, measured on an NVIDIA RTX A6000 GPU), and parameter count (in millions), respectively.