When transferring manipulability across systems with different sizes and kinematic structures, matching absolute ellipsoid scale may be unnecessary when the goal is to reproduce orientation and semi-axis length ratios. Full-matrix tracking, however, penalizes both shape and absolute-scale differences, even when only shape matching is required. We therefore propose a scale-invariant manipulability shape-tracking method that treats matrices differing only by a positive scalar factor as equivalent and uses their unit-determinant representatives. We derive the differential of the unit-determinant shape representative and an orthonormal coordinate representation of the tangent tracking residual under the affine-invariant Riemannian metric (AIRM). The resulting scale-invariant objective is integrated with position and end-effector direction tasks in a constrained joint-velocity quadratic program. Simulations with four heterogeneous robots evaluate robot-to-robot and human-to-robot transfer. On three followers, the proposed method achieves endpoint shape distances of 9.30 x 10^-5 without scale tuning. With robot-specific target scales tuned during motion, the Full method retains endpoint axis-ratio errors of 0.19-0.31 on KR500 and UR20. For human reaching with concurrent tasks, the proposed method yields dual force shapes elongated along X like the human target on all four robots, with endpoint position errors of 2.4-5.6% of reference arm length versus up to 75% for the Full method tracking the original human ellipsoid.
Figures & tables
Fig. 1: Manipulability ellipsoid (ME) shape transfer from robot or human motion to heterogeneous robots. \scriptsize1⃝ Reference and current manipulability matrices are normalized to unit-determinant shapes, preserving ME orientation and all semi-axis length ratios. \scriptsize2⃝ The Shape method follows the reference shape trajectory without requiring absolute-scale matching. Blue dashed ellipses denote reference MEs; green, red, and orange solid ellipses denote the MEs of Gen3, KR500, and UR20, respectively.
Fig. 2: Manipulability transfer from FR3 to Gen3, KR500, and UR20. (a) Robot motions: FR3 source (blue frame) and followers using the Full (teal) and Shape (orange) methods. (b) Actual (red solid) and desired (blue dashed) XZ ellipses centered at the corresponding times, displayed as M for the Full method and M for the Shape method.
Fig. 3: Human-to-robot transfer of wrist position and manipulability shape from a measured right-arm reach. Four heterogeneous manipulators use the Shape method to track the transferred position and velocity manipulability shape with a fixed forward end-effector direction objective. Ellipsoids show the final human target (blue) and robot (red) dual force shapes.
Phase means
Full
Shape
Robot
Phase [s]
dAI
ds
dρ
dAI
ds
dρ
Gen3
[0,3)
0.0348
0.0344
0.0055
0.0651
0.0344
0.0421
[3,8)
0.2210
0.2210
0.0033
0.2234
0.2210
0.0176
[8,11]
0.0409
0.0408
0.0021
0.0471
0.0408
0.0101
KR500
[0,3)
2.7018
0.8115
2.5136
3.0878
0.0243
3.0873
TABLE I: Tracking errors
Endpoint ( t=11 s)
Follower
Multiplier c
∣r−rd∣
θ[deg]
ds
Gen3
1.189
0.0002
0.0008
0.0001
11.31
1.0687
0.0005
0.4352
5.657
0.9562
0.0005
0.3826
KR500
1.189
4.5546
0.3019
1.0694
11.31
0.1905
0.0004
0.0684
TABLE II: Target-scale sensitivity of the Full method
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dairu Liu, Zekun Qi, Jiayu Zeng +11
1Nankai University · 3Galbot · 2Tsinghua University +3
Human demonstrations provide strong priors for robot manipulation, yet it is non-trivial to transfer them to execute on real robots due to the kinematic gap. In dexterous manipulation, it remains challenging to track long-horizon, contact-rich sequences even in simulators: a reference-tracking policy must keep objects on their target trajectories while preserving demonstrated joint motion and contact timing. Existing approaches often rely on hand-crafted reward tuning that require per-sequence tuning and break under limited interaction budgets. We introduce ConTrack, a reinforcement learning (RL) framework that scales with tracking data. ConTrack treats object tracking as a constraint and allocates remaining control authority to motion fidelity, which allows it to adapt task--style trade-offs online using a dual-variable update. In addition, ConTrack also stabilizes long-horizon learning with an adaptive mid-trajectory reset library that reuses policy-reachable simulator states. Our qualitative and quantitative results in simulation tracking and real robot demonstrate that ConTrack improves success and object pose accuracy significantly over prior arts while preserving joint and contact fidelity. Website: https://www.lyt0112.com/projects/ConTrack.
A humanoid with 5-DoF arms cannot track generic bimanual end-effector trajectories with its arms alone; pelvis and waist motion must be recruited, but which motion, and when, is not uniquely determined. On a Unitree R1 in fixed double support, the set of dynamically valid recruitment strategies (pelvis pose and waist trajectories) for a task is a diverse continuous manifold, and a deterministic regressor trained on it mode-averages into strategies valid only 35% of the time, against 52% for a conditional variational autoencoder (CVAE) and 82% for the best of 16 CVAE samples. We introduce AMBIT: the CVAE proposes strategies from a preview of the commanded trajectory, a non-learned selector filters, ranks and verifies them, and a receding-horizon loop commits to one with hysteresis. The committed strategy is the reference of the same whole-body differential-IK QP a reactive tracker runs, which keeps authority over residual error. On 160 held-out episodes that admit a valid strategy, in full MuJoCo dynamics under a torque controller, AMBIT reaches 85% success at a 3 cm/15 deg tolerance against 74% for the tracker (disjoint confidence intervals) and recruits the body before the arms saturate in 48% of episodes against 35%. Because diversity is preserved, constraints unknown at training time are enforced by selection alone: under five zero-shot shifts AMBIT beats the warm-started tracker on every shift and matches a test-time re-optimisation baseline 17x more expensive. On a Unitree G1, with hyperparameters unchanged, the protocol reproduces the structure of the valid set and widens the gap over the tracker to 0.85 against 0.53. Five selected strategies execute on the externally supported physical R1, distinct in pelvis excursion and tracking the planned end-effector motion to a median of 11 mm by encoder forward kinematics, which establishes kinematic realisability, not balance.