cs.ROSep 30, 2026

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

Authors: Wenhao Li, Xiu Su, Yu Han, Yichao Cao, Shan You, Chang Xu

Organizations: University of Sydney · Central South University · University of California, San Diego · Ace Robotics

Abstract

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6×\times the success rate of PI0.5 in real-world dynamic settings and about 5×\times on DynBench. We will open-source all the code and weights.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Generalizable Robotic Manipulation in Dynamic Environments

    Mar 16, 2026Heng Fang, Shangru Li, Shuhan Wang +3Robotic ManipulationDynamic Environments

  2. Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation

    Jun 1, 2026Shahram Najam Syed, Arthur Jakobsson, Haoran Hao +1Future Video PredictionLatent World Models