CALM: Current Aligned Link Manipulation for Single Arm Oversized Object Lifting
Authors: Jun Hu, Sihan Chen, Kosta Jovanovic, David Navarro-Alarcon, Xueqian Wang, Jia Pan, Peng Zhou
Organizations: School of Advanced Engineering, Great Bay University, Dongguan, Guangdong, China · Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR, China · School of Electrical Engineering, University of Belgrade, Belgrade, Serbia · Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong SAR, China
Most robots manipulate objects solely with their end effectors, whereas humans flexibly leverage different body parts, such as the forearm and elbow, especially when handling oversized objects. Learning such whole-arm manipulation is chal-lenging due to long-horizon sparse rewards, limited contact sens-ing, and the sim-to-real gap in contact and actuator dynamics. To address these challenges, we propose Current-Aligned Link Manipulation, a framework for learning long-horizon contact-rich manipulation using motor current as joint load related feedback. Three stage-specific policies first learn repositioning, grasping, and lifting using privileged simulation information, and a stage router sequences them to generate complete task demonstrations. For sim-to-real transfer, a causal current mapper predicts physical motor current from simulated joint histories, aligning the actuator current observation between simulation and hardware. A unified student policy then learns from these demonstrations using only deployable sensor observations and is further refined with DAgger. The task policies are trained entirely in simulation, and the final student is deployed on hardware. Experiments demonstrate 76.2% (762/1000 trials) complete-task success in simulation and 73.3% success (22/30 trials) on the physical robot for sequential oversized-object lifting.
Figures & tables
Fig. 1: Motivation and task overview. (a) When an object’s width exceeds the functional grasp aperture of one hand, (b) a person can recruit the whole arm to support and carry it. Analogously, (c) when the object’s width exceeds the aperture of the robot’s end effector gripper, we investigate whether a single rigid manipulator can use intermediate arm links to (d) reposition the object, (e) establish an opposed grasp, and (f) lift it.
Fig. 2: Overview of CALM. A causal Transformer maps simulated joint histories to physical motor current, while three privileged teachers learn repositioning, grasping, and lifting. During complete task rollouts, privileged simulation state selects the active teacher to produce complete trajectories under domain randomization. These trajectories supervise one stage aware student, which is refined with DAgger and deployed with deployable observations only.
Reward: (1) contact at designated links; (2) centered opposed grasp; (3) stable hold. Penalty: (1) missing or invalid contact; (2) box motion; (3) grip loss; (4) table contact; (5) motion and force variation after latching.
Lift πL
Reward: (1) grasp confirmation and maintenance; (2) lift progress and height; (3) settling and holding. Penalty: (1) weak or lost grasp; (2) premature or stalled lifting; (3) lateral motion; (4) wrist slip; (5) large actions before confirmation.
TABLE II: Principal rewards and penalties specific to each stage.
Question summary
Experiments providing evidence
Current alignment
Secs. IV-B and IV-E : current prediction accuracy and 10 hardware trials without current alignment
Student distillation
Secs. IV-C and IV-D : controlled comparison of the full student policy and its ablated variants
Physical CALM
Sec. IV-E : 30 hardware trials of the final CALM policy and a current alignment ablation comprising 10 additional trials, with a new random box placement before each trial
TABLE III: Research questions and corresponding experiments.
Fig. 3: Experimental setups. (a) Isaac Lab simulation environment. (b) Physical robot setup, where two cameras detect ArUco markers attached to the box to track its position.
Fig. 4: Offline current alignment. (a) Raw simulated joint effort (right axes), recorded physical motor current, and current predictions (left axes) over the first eligible continuous 15 s interval. (b) MAE for each joint and the mean across joints over all 19,213 held out samples. The simulated joint effort baseline is converted to amperes using a separate linear scaling for each joint fitted on the training data.
Fig. 5: Teacher routing and teacher to student transfer in simulation. (a) Routed teacher stage progression. Each stage reports routed reach followed by independently evaluated teacher success, and each failure branch terminates at the first unreached stage. (b) Ablation of stage supervision and DAgger. Cells report unconditional reposition completion, conditional R→G and G→L success, and overall task success over 1,000 episodes per policy.
Fig. 6: Physical CALM evaluation under random initial box placements. From left to right, each row shows the initial state, box repositioning, opposed link grasping, and the lifting outcome. (a) and (b) Successful trials with the small box. (c) Successful trial with the large box. (d) Failed trial in which the large box is dropped after grasping.
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
A key challenge in contact-rich dexterous manipulation is the need to jointly reason over global geometry and nonsmooth contact dynamics. End-to-end policies bypass this complexity, but often require large amounts of data and transfer poorly from simulation to reality. We address the limitations with a simple insight: dexterous manipulation is inherently hierarchical--at a high level, a robot decides where to touch (geometry); at a low level it determines how to move the object through contact dynamics. Building on this insight, we propose a hierarchical RL--MPC framework in which a high-level reinforcement learning (RL) policy predicts a contact intention, a novel object-centric interface that specifies (i) an object-surface contact location and (ii) a post-contact object subgoal pose. Conditioned on the contact intention, a low-level contact-implicit model predictive control (MPC) optimizes local contact modes and real-time (re)plans through contact dynamics to generate robot actions that robustly move the object toward each subgoal. We evaluate the framework on non-prehensile tasks, including geometry-generalized pushing across diverse object shapes, pivoting/flipping-based object reorientation, and environment-assisted object repositioning. It achieves high success rate with substantially reduced data (10 times less than end-to-end baselines), highly robust performance, and zero-shot sim-to-real transfer.
Zhixian Xie, Yu Xiang, Michael Posa +1
Arizona State University · University of Texas at Dallas · University of Pennsylvania