DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation
Authors: Chengxi Li, Yan Di, Yingyue Li, Ruida Zhang, Mingyang Li, Xiangyang Ji
Organizations: Department of Automation, Tsinghua University, China · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China · Beijing Institute of Control Engineering, China
Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46× faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.
Figures & tables
Fig. 1: DualManip decouples semantic reasoning from geometric grasp adaptation. As the toy snake deforms, the geometric path updates the grasp while preserving the task intent established by the semantic path. In contrast, ReKep and OmniManip fail to maintain a valid task-consistent grasp under the same non-rigid deformation.
Fig. 2: Overview of DualManip. Given instruction and RGB-D observation, the semantic path establishes a task-conditioned plan, from which the information interaction module generates the initial grasp. During execution, the geometric path continuously updates the grasp from live RGB-D observations, executing feasible updates and falling back otherwise. For subsequent manipulation stages, task-conditioned spatial relations are formulated as constraints and optimized into executable poses.
Fig. 3: Illustration of feasibility checking. The green gripper denotes the updated grasp, green points the matched grasp contacts, and red markers infeasible interactions. (1) Collision-prone and (2) unintended grasps are rejected before execution, triggering semantic replanning.
Category
Task
ReKep
OmniManip
CLEA
DualManip (Ours)
General
Snake Placement
8/15
12/15
11/15
12/15
Toy Positioning
10/15
11/15
12/15
12/15
Block Placement
9/15
12/15
11/15
11/15
Mean
60.0%
77.8%
75.6%
77.8%
Assembly
RAM Insertion
1/15
0/15
0/15
4/15
GPU Insertion
4/15
6/15
5/15
7/15
TABLE I: Quantitative results under the static setting. Entries denote successful trials, while Mean reports average success rate (%).
Fig. 4: Representative execution results for the high-precision assembly tasks under the static setting, highlighting DualManip’s ability to perform precise task-conditioned insertion and placement.
Setting
Task
ReKep
OmniManip
CLEA
Open-loop
DualManip (Ours)
Single-change
Snake Placement
5/15
2/15
9/15
0/15
10/15
Toy Positioning
7/15
1/15
10/15
0/15
9/15
Block Placement
7/15
9/15
10/15
0/15
10/15
Mean
42.2%
26.7%
64.4%
0.0%
64.4%
Continuous
Snake Placement
3/15
1/15
0/15
0/15
8/15
Toy Positioning
4/15
0/15
0/15
0/15
7/15
TABLE II: Quantitative results under dynamic settings. Entries denote successful trials, while Mean reports average success rate (%).
Fig. 5: Representative executions under the single-change setting. Top: toy-positioning under articulated reconfiguration. Bottom: block-placement under rigid displacement.
Fig. 6: Qualitative execution on snake-placement under continuous scene variations. DualManip continuously updates grasp poses in real time to accommodate object motion and deformation while maintaining valid task intent.
Method
ReKep
OmniManip
CLEA
DualManip (Ours)
GPT
Qwen
Latency
77.2 ms
231.3 ms
6.376 s
4.329 s
138.7 ms
TABLE III: Quantitative comparison of average adaptation latency under a single scene change.
Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul +2
Khalifa University, Abu Dhabi, United Arab Emirates. · Sereact GmbH, Stuttgart, Germany.
Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The second is state aliasing, where visually similar observations from different points in a task require different actions. We argue that these failures persist regardless of model scale and inference latency, showing that the bottleneck is missing temporal context rather than model capacity. Based on this insight, we propose TEMPO, which augments a pretrained VLA with two temporal inputs: a motion summary extracted from a frozen video foundation model to resolve motion ambiguity and a compact proprioceptive history to resolve state aliasing. TEMPO requires no modification to the backbone and adds minimal compute overhead at training or deployment. Across four dynamic manipulation tasks, it improves Bottle Handover success from 44% to 74% and is the only method that solves state aliasing. Probing and ablation studies confirm that each temporal signal independently addresses its corresponding failure. We further release TEMPO-Bench, a benchmark of over 50k annotated frames for evaluating motion-aware robot perception in both regression and multiple-choice formats. Project Website: https://tempo-robot.github.io/
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception-execution mismatch: object states continue to evolve during inference, making actions predicted from past observations stale at execution time. We present DynamicVLA, a latency-aware VLA model for dynamic object manipulation. It combines a compact 0.4B architecture and convolutional vision encoder for efficient multimodal inference with a continuous inference schedule that overlaps reasoning and execution for non-blocking control. Latent-aware Action Streaming then discards latency-invalid action prefixes and executes only the temporally valid suffix of each predicted chunk, preserving action-time alignment under dynamic object motion. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations in simulation and on real robots show that DynamicVLA improves dynamic manipulation success under changing object motion, perception-heavy instructions, and unseen motion patterns.