Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without distorting the trajectory shape. Another key limitation is that the resulting policies often lack precise object-level 3D geometry awareness, limiting object grounding and object shape awareness critical for precise manipulation. To bridge these gaps, we propose VidAct, an efficient video-to-robot framework that learns object-centric, 3D-aware manipulation policies from a single monocular video per task and enables zero-shot real-world deployment. VidAct consists of three key components. First, VidAct reconstructs object meshes and motion from arbitrary demo videos and canonicalizes the motion in the static object frame, avoiding embodiment-specific retargeting and accommodating diverse camera viewpoints. Second, VidAct employ residual trajectory transfer for adapting the reconstructed motion to novel object configurations while preserving its motion shape. Finally, as the key policy-learning component, VidAct predicts simulation-provided privileged complete-object point clouds at each frame as an auxiliary task while retaining RGB-only deployment, providing dense object-centric supervision over both object pose and 3D geometry. Experiments on human, robot, generated, and internet videos demonstrate broad video applicability and zero-shot deployment. Per-frame complete-object 3D supervision improves policy generalization and sim-to-real success, while residual trajectory transfer enables reliable trajectory adaptation with better shape preservation.
Figures & tables
Method
Single RGB Video
Embodiment Agnostic RGB Video
Retargeting Free
RL Free
DemoGen [ 7 ]
✗
✗
✓
✓
R2R2R [ 6 ]
✗
✗
✗
✓
X-Sim [ 3 ]
✗
✗
✓
✗
RigVid [ 5 ]
✗
✗
✓
✓
DexImit [ 4 ]
✓
✗
✗
✓
VidAct (Ours)
✓
✓
✓
✓
TABLE I : Comparison with existing data generation methods.
Fig. 2 : The pipeline of VidAct. VidAct reconstructs object meshes and trajectories from videos and transforms the trajectories into the object frame. Residual trajectory transfer efficiently augment the reconstructed trajectories to diverse object configurations. After selecting a grasp pose of gripper or Dexhand, VidAct derives the corresponding end-effector poses, rolls out the trajectories in simulation to generate training data, and trains an object-centric 3D-aware policy for zero-shot deployment.
Fig. 3 : Residual trajectory transfer for trajectory augmentation. Left: illustration of residual trajectory transfer; right: sampled novel object poses.
Fig. 4 : Model architecture of object-centric 3D predict policy. We add learnable query and auxiliary head to ACT [ 9 ] .
Fig. 5 : Representative videos processed by VidAct.
Fig. 6 : Representative videos and the corresponding simulation tasks processed by VidAct. Top: Videos; bottom: Simulation. The video sources are generated from Wan2.2 [ 36 ] and from UMI [ 39 ] , DreamDojo [ 38 ] and BridgeData V2 [ 37 ] .
Source Video Type
Overall
Mask gen.
Mesh recon.
Pose tracking
Self-collected
4/4
4/4
4/4
4/4
Internet Human
14/20
19/20
16/19
14/16
Internet Robot
15/22
20/22
18/20
15/18
Generated
9/20
17/20
14/17
9/14
TABLE II : Overall conversion rates across video sources and pipeline reliability across video sources.
Fig. 7 : Zero-shot real-robot rollout visualizations for four tasks. From top to bottom, each row corresponds to Cup Stacking, Ketchup Pouring, Flower Insertion and Bag Hanging.
Fig. 8 : Performance of standard visuomotor policy ACT [ 9 ] (2D) and our object-centric 3D prediction policy in simulation and the real world across varying sizes of training data. All policies are trained on simulated data without background randomization.
Fig. 9 : Visualization of real-world point cloud predictions projected onto the images.
Fig. 10 : Generalization setup with diverse backgrounds and varied table textures and colors in real world.
Simulation
Real World
Task
ACT
ACT_Pose
ACT_3D
ACT
ACT_Pose
ACT_3D
Cup Stacking
88%
89.5%
98%
17/30
18/30
21/30
Ketchup Pouring
77.5%
79%
93.5%
10/30
11/30
15/30
Flower Insertion
60%
58.5%
71%
6/30
6/30
10/30
Bag Hanging
63.5%
65%
76.5%
8/30
9/30
12/30
TABLE III : Generalization comparison of ACT, ACT_pose and ACT_3D. The best result is bold .
Task
Interpolation
Replay
Residual Traj. Transfer (Ours)
Trajectory Generation Success Rate
Cup Stacking
69.1%
70.6%
96.4%
Ketchup Pouring
83.6%
39.7%
95.8%
Flower Insertion
24.8%
33.2%
85.4%
Bag Hanging
17.0%
81.1%
98.2%
Average
48.63%
56.15%
93.95%
TABLE IV : Comparison of residual trajectory transfer with trajectory augmentation baselines. The best result in each row is bold .
Achieving autonomous robotic dexterous manipulation requires precise, human-like action sequences at scale. As a scalable supplement to costly teleoperation data, extracting trajectories with both visual fidelity and physical plausibility from monocular videos represents a promising frontier in embodied AI. To this end, we introduce V2P-Manip, an efficient framework designed to learn dexterous manipulation policies directly from human demonstration videos. We establish an efficient, integrated pipeline encompassing 3D asset acquisition, trajectory estimation, and dexterous policy learning. To bridge the gap between visual perception and physical constraints, we introduce a two-stage refinement process to enforce spatial alignment and physical consistency. Evaluations on the TACO and OakInk benchmarks demonstrate that our approach significantly outperforms previous methods in pose accuracy, adaptability to unstructured environments, and training efficiency. Ultimately, experimental results confirm an average success rate of over 75% across multiple synthetic manipulation tasks and validate the adaptability of the extracted manipulation priors across diverse dexterous hand embodiments.
Kaihan Chen, Yanming Shao, Haifeng Ji +2
Zhejiang University, Hangzhou, China · Shanghai Jiao Tong University, Shanghai, China · Shanghai AI Laboratory, Shanghai, China
Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to capture complex geometric constraints or generate goal-state observations to capture semantic and geometric knowledge, but fail to explicitly couple object transformation with action prediction, resulting in errors due to generative noise. To address these limitations, we propose Imagine2Act, a 3D imitation-learning framework that incorporates semantic and geometric constraints of objects into policy learning to tackle high-precision manipulation tasks. We first generate imagined goal images conditioned on language instructions and reconstruct corresponding 3D point clouds to provide robust semantic and geometric priors. These imagined goal point clouds serve as additional inputs to the policy model, while an object-action consistency strategy with soft pose supervision explicitly aligns predicted end-effector motion with generated object transformation. This design enables Imagine2Act to reason about semantic and geometric relationships between objects and predict accurate actions across diverse tasks. Experiments in both simulation and the real world demonstrate that Imagine2Act outperforms previous state-of-the-art policies. More visualizations can be found at https://sites.google.com/view/imagine2act.
Liang Heng, Jiadong Xu, Yiwen Wang +6
CFCS, School of Computer Science, Peking University · PrimeBot · AGIBOT
Human manipulation videos provide rich motion and interaction cues for acquiring robot skills without robot demonstrations. Video generation models synthesize such demonstrations from an initial scene image and task instruction, avoiding the need to record demonstrations for each task. However, the recovered motion captures only one scene-specific realization, leaving task structure, geometric relations, and constraints implicit. We present V2-STRep, a zero-shot framework that converts generated video motion into reusable robot skills through VLM-grounded structured task representations. The representation specifies motion phases, references, and task-relevant constraints, with targets described by minimal geometric structures: points, point-normals, axes, planes, and full 6D poses. VLM-provided 2D image-space cues are lifted into 3D using RGB-D observations to reconstruct task geometry and candidate grasp poses. Geometry-specific rules transfer motion to new scenes, while task-constrained trajectory optimization couples grasp selection with complete robot motion planning. It preserves task requirements while using remaining rotational freedom to accommodate joint limits. Updating deployment grounding and constraints enables reuse under new compatible instructions without generating another video. Experiments on six real-world manipulation tasks demonstrate improved execution success over baselines, reliable cross-scene transfer of successfully acquired skills, and adaptation to changed deployment instructions.
Yexin Hu, Dongheui Lee
Autonomous Systems Lab, Technische Universität Wien (TU Wien), Austria · Institute of Robotics and Mechatronics, German Aerospace Center (DLR), Germany