Dexterous teleoperation requires reliable human-hand state estimations. However, common low-cost motion-capture gloves and markerless trackers often exhibit biases that vary across users, glove fit, and recording sessions, degrading retargeting and demonstration quality. We present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses. Instead of optimizing a separate model for every operator or session, YOCO conditions a calibration HyperNet on the paired examples and predicts LoRA-style updates for a frozen MANO hand-estimation module, turning per-user calibration into a lightweight feed-forward adaptation step while preserving the geometric prior of MANO and the efficiency of a compact estimator. We train YOCO with synthetic drift augmentations on InterHand2.6M and evaluate on augmented InterHand sequences, offline real glove data, and dexterous teleoperation tasks. Across these settings, YOCO improves calibration efficiency, hand-state estimation quality and teleoperation performance compared with uncalibrated input and standard calibration baselines.
Figures & tables
Figure 1: Overview of YOCO . Low-cost mocap streams suffer from user-, device-, and session-specific bias. Given a few paired calibration poses, YOCO predicts a calibration update that corrects future mocap poses and improves downstream dexterous teleoperation.
Figure 2: YOCO model architecture. Given a calibration set S={(xin,xic)}i=1K of K paired noisy and target poses (left), the calibration HyperNet hψ (middle) encodes each pair, aggregates them through a transformer-based calibration mapper with one CLS token per LoRA matrix, and decodes the tokens into LoRA updates ΔW . These updates are injected once per session into the frozen MANO estimator gϕ (right) to form gϕ+ΔW , which maps each raw mocap query xqn at frame rate to calibrated MANO pose and 3D joints.
Method
MPJPE ↓
PA-MPJPE ↓
Tip-PA ↓
AUCJ ↑
Pinch ↓
Calib. s ↓
(mm)
(mm)
(mm)
(mm)
(s)
No calibration
21.52
16.55
37.15
0.611
48.49
0.00
Ridge
17.62
13.00
28.28
0.676
45.35
1.6
Last-layer-Finetune
16.62
12.01
25.92
0.692
39.18
45.6
Full-Finetune
18.11
12.77
29.21
0.668
37.11
95.4
LoRA-Finetune
16.38
11.67
25.33
0.694
37.28
89.4
Table 1: Calibration results on the augmented InterHand2.6M evaluation setting. Calib. s reports the total adaptation time in seconds for whole validation sets on one H100.
Figure 3: PA-MPJPE versus calibration-pair count, including YOCO followed by test-time fine-tuning (YOCO-FT). The dotted line marks no calibration.
Figure 4: Goal states for the downstream dexterous manipulation tasks: twisting and opening a bottle cap, pinch-grasping earbuds, three-finger glue grasping, and tissue-box opening.
Method
MPJPE ↓
PA-MPJPE ↓
Tip-PA ↓
AUCJ ↑
Pinch ↓
Calib. ↓
(mm)
(mm)
(mm)
(mm)
(s)
No calibration
20.49
14.63
36.22
0.610
51.39
0.00
Ridge
24.71
19.21
45.50
0.573
57.54
1.49
Last-layer-Finetune
18.86
14.67
33.88
0.643
43.63
5.04
Full-Finetune
20.00
15.73
36.67
0.630
49.89
13.04
LoRA-Finetune
18.72
13.94
33.23
0.648
54.01
14.57
Table 2: Real-world mocap calibration results. Metrics are averaged over 3 seeds. Calib. s reports the total conditioning or adaptation time in seconds for whole datasets on one H100. Corresponding user-level std is provided in Appendix C.5 .
Method
Bottle Cap [-0.3ex]Opening
Earbud Pinch Grasping
Glue Three-Finger [-0.3ex]Grasping
Tissue Box Opening
Overall
Success ↑
Time (s) ↓
Success ↑
Time (s) ↓
Success ↑
Time (s) ↓
Success ↑
Time (s) ↓
Success Rate ↑
Time (s) ↓
No calibration
9/10
35.6
0/10
–
5/10
7.0
0/10
–
35%
24.4
LoRA-Finetune
9/10
34.1
1/10
6.0
8/10
6.6
8/10
13.3
65%
17.9
YOCO (ours)
10/10
27.3
8/10
4.2
8/10
8.9
10/10
17.4
90%
14.9
Table 3: Downstream dexterous teleoperation results. We evaluate 4 tasks and count each task success over 10 trials and average completion time over successful trials. Failed trials are capped at one minute and excluded from task-specific average completion time, while each failed trial is penalized by assigning it the longest observed trial time in the overall average time. More stats are provided in Appendix C.8 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Real glove (AVP reference)
Augmented InterHand2.6M
Method
PA-MPJPE ↓
AUCJ ↑
PA-MPJPE ↓
AUCJ ↑
LoRA-Finetune
13.94
0.648
11.67
0.694
Set-Cond. Residual
13.74
0.645
12.66
0.664
Side-Tuning [ 6 ]
13.48
0.641
15.35
0.618
Meta-LoRA
13.51
0.633
13.94
0.639
Meta-LoRA+FT
14.17
0.664
11.41
0.702
Appendix
Table 4: Learned-prior controls and additional calibration baselines at K=4 . PA-MPJPE is in millimeters; AUCJ is the normalized joint-PCK area.
Figure 5: Preset pose library used to obtain deployable calibration targets. The four prompts are index pinch, fully open, claw / half close, and four fingers straight with the thumb near the palm.
Method
Timing Source
CPU calibration time (s)
YOCO (ours)
HyperNet Forward
13.16
Ridge
Closed Form Solver
0.37
Last-layer-Finetune
Gradient Descent
34.47
LoRA-Finetune
Gradient Descent
59.58
Full-Finetune
Gradient Descent
223.25
Appendix
Table 5: Average CPU calibration time required before real deployment.
Method
MPJPE ↓
PA-MPJPE ↓
Tip-PA ↓
AUCJ ↑
Pinch ↓
Calib. s ↓
(mm)
(mm)
(mm)
(mm)
(s)
No calibration
1.14
0.76
3.18
0.021
7.01
0.00
Ridge
4.32
2.95
9.12
0.053
12.53
0.68
Last-layer-Finetune
0.87
1.01
2.47
0.017
6.23
0.55
Full-Finetune
1.80
1.52
4.79
0.030
7.65
0.75
LoRA-Finetune
1.22
1.06
3.73
0.023
9.50
0.80
Appendix
Table 6: Per-user standard deviation of the real-world mocap calibration results.
Method
K
MPJPE ↓
PA-MPJPE ↓
Tip-PA ↓
AUCJ ↑
Pinch L2 ↓
(mm)
(mm)
(mm)
(mm)
YOCO (ours)
1
16.07
10.94
24.40
0.694
38.04
2
16.10
10.86
24.30
0.694
36.66
4
15.64
10.45
23.29
0.702
34.90
8
15.21
10.19
22.26
0.709
33.87
16
15.10
10.14
21.94
0.712
33.70
Appendix
Table 7: Detailed K-sweep results on the augmented InterHand2.6M evaluation setting. All rows use the overall split aggregated over 3 seeds and 20 augmented units.
Method
Bottle Cap Opening
Earbud Pinch Grasping
3-Finger Glue Grasping
Tissue Box Opening
Avg. Time (s) ↓
Avg. Time (s) ↓
Avg. Time (s) ↓
Avg. Time (s) ↓
No calibration
37.9
7.0
10.5
42.0
LoRA-Finetune
37.5
6.9
8.1
19.0
YOCO (ours)
27.3
4.8
9.9
17.4
Longest trial
59.0
7.0
14.0
42.0
Appendix
Table 8: Detailed downstream teleoperation completion-time statistics. Values are average completion time in seconds over 10 trials; lower is better. The longest-trial row reports the longest observed trial time for each task.
Component
Value
CPU
2 × Intel Xeon Platinum 8468V
RAM
2.0 TiB
GPU
1 × NVIDIA H100 80 GB HBM3
Appendix
Table 9: Device and compute platform used for training and timing experiments.
Teleoperation is a key interface for controlling dexterous robotic hands and collecting demonstrations for imitation learning. Its effectiveness largely depends on kinematic retargeting, which maps operator hand motions to feasible and intuitive robot hand motions. Existing methods often require hand-crafted objectives, precise calibration, or global shape matching between human and robot hand spaces, making them sensitive to hand-specific tuning and less reliable across different dexterous hands. We propose AnyDexRT, a calibration-free retargeting method for intuitive dexterous teleoperation across human-like dexterous hands. AnyDexRT combines self-supervised fingertip correspondence learning with few-shot human guidance to anchor the mapping in task-relevant regions, and further refines pinch-related poses using a contact classifier. Experiments on diverse dexterous hands and real-world teleoperation tasks show that AnyDexRT improves retargeting quality, reduces manual tuning, and provides more intuitive and efficient control than prior retargeting methods. Project website: https://chenxi-wang.github.io/projects/anydexrt
Chenxi Wang, Ying Feng, Hongjie Fang +4
1Noematrix · 2Shanghai Jiao Tong University · 3Shanghai Innovation Institute
This work presents an RGB-D imaging-based approach to marker-free hand-eye calibration using a novel implementation of the iterative closest point (ICP) algorithm with a robust point-to-plane (PTP) objective formulated on a Lie algebra. Its applicability is demonstrated through comprehensive experiments using three well known serial manipulators and two RGB-D cameras. With only three randomly chosen robot configurations, our approach achieves approximately 90% successful calibrations, demonstrating 2-3x higher convergence rates to the global optimum compared to both marker-based and marker-free baselines. We also report 2 orders of magnitude faster convergence time (0.8 +/- 0.4 s) for 9 robot configurations over other marker-free methods. Our method exhibits significantly improved accuracy (5 mm in task space) over classical approaches (7 mm in task space) whilst being marker-free. The benchmarking dataset and code are open sourced under Apache 2.0 License, and a ROS 2 integration with robot abstraction is provided to facilitate deployment.
Martin Huber, Huanyu Tian, Christopher E. Mower +4
School of Biomedical Engineering & Imaging Sciences, King’s College London, London, UK · Noah’s Ark Lab, Huawei, London, UK · independent contributor
Assembly, wear, and component replacement perturb the sensor extrinsics and joint zeros encoded by a humanoid CAD model. Existing procedures calibrate one sensor pair or require external fiducials. Using only robot-native motion and onboard sensing, we present OmniCalib, a target-free workflow that calibrates the full upper limbs---all 14 arm joint zeros and the extrinsics of both wrist and chest cameras---as well as lower limbs and the multi-camera head rig. Each module matches a robot-native task to a parameter block, checks observability, and writes only supported corrections to the CAD model. Our depth ICP method recovers all 14 arm joint zeros and calibrates all RGB-D camera extrinsics without any calibration target. Relative to CAD, the estimated extrinsic corrections are 10.56 mm and 1.74 degrees for the left wrist, 6.33 mm and 1.25 degrees for the right wrist, and 9.81 mm and 0.929 degrees for the chest RGB-D camera. ICP point-to-plane residual is 2.09 mm. On the same injected offsets, ICP and ArUco recover all 14 joint zeros below the 0.1-degree encoder-resolution reference. On an AGIBOT A3 Ultra humanoid, four static double-support stances recover all 12 lower-limb joint-zero offsets injected with an RMS error of 0.063 degrees. The head module combines multi-camera visual odometry with legged odometry and dynamic compensation through the live ROS transform tree. Using only planar walking, it attains a mean SO(3) error of 1.061 degrees across three sequences. The best sequence reaches 0.775 degrees, competitive with iKalibr at 0.902 degrees from rich 6-DOF excitation. Rig-relative angles repeat within 0.140 degrees. Injection recovery and held-out tests validate each observable block.