Fast Non-Parametric Heteroscedastic Imitation Learning With Geometric Priors
Authors: Maximilian Mühlbauer, Arne Sachtler, Markus Knauer, Cem Küçükgenç, Yanlong Huang, Alin Albu-Schäffer, João Silvério
Organizations: Department of Computer Engineering, Technical University of Munich, Friedrich-Ludwig-Bauer-Str. 3, Garching, Germany. · German Aerospace Center (DLR), Robotics and Mechatronics Center (RMC), Münchener Str. 20, 82234 Weßling, Germany. · School of Computer Science, University of Leeds, Leeds, UK.
When learning probabilistic policies from human demonstrations, data-efficient learning and fast adaptations to new scenarios are key requirements. One popular way to achieve intuitive and reliable adaptations is through non-parametric, typically kernel-based, methods. However, existing solutions either fail to account for the geometry of manifolds common in robotics, limiting data efficiency, or, when geometry-aware, provide unreliable uncertainty estimates or require retraining to adapt. We propose a non-parametric approach leveraging geometric priors in scenarios of data scarcity and heteroscedastic uncertainties for probabilistic modeling. We utilize the method to formulate policies based on time or robot state, where non-separable diagonal kernels allow capturing uncertainty relations between degrees of freedom for same-sized in- and outputs. Fast updates, requiring less than 3 ms for a trajectory involving both position and orientation are possible through an optimized formulation. Our approach supports both manifold-valued input and manifold-valued output with large orientation changes. Using task parameterization, adaptation to different object poses is easily possible. We evaluate the approach on a set of toy examples and on real robot manipulation tasks both in autonomous execution and in shared control scenarios.
Figures & tables
Fig. 2 : Manifolds used in this work, proposed by [ 21 ] . Cylindrical ( M2 ) or spherical ( M3 ) coordinates allow expressing some tasks more data-efficient compared to Cartesian ( M1 ) coordinates. Coordinate systems in each image depict the unit quaternion (0,0,0,1)⊤ for different positions on the manifold.
Fig. 3 : Left : Pose data 1 reproduced by a Riemannian KMP ( III-A ) 2 with correct rotations. In contrast, the STS KMP [ 9 ] 3 suffers from incorrect rotations. Right : Adding a via point V 0.1m above the trajectory with a relative rotation of −0.35rad about the y axis and a null space action N . Via point and nullspace target are visualized by a larger coordinate system.
Fig. 4 : Plots for the letter “B” of the handwriting dataset [ 10 ] projected to S2 . We fit the GMM using 10 Gaussians and use 20 reference points for KMP ; all these Gaussians are depicted with their mean and covariance ellipsoid.
Method
Position Difference (mm)
Rotation Difference (mrad)
GMR ( STS )
2.47±1.02
3139.36±11.47
KMP ( STS ) [ 9 ]
2.47±1.03
269.36±646.48
Riemannian GMR [ 4 ]
2.47±1.09
15.27±8.81
Riemannian KMP (ours)
2.46±1.12
15.15±9.05
Nadaraya-Watson [ 17 ]
3.05±1.22
15.29±9.15
TABLE I : Differences between the prediction of different methods and ground truth for the spiral trajectory ( Fig. 3 ).
Kernel
S1
S1×R2
S2
SO(3)
R3×SO(3)
RBF
0.01
0.65
0.27
0.63
1.1
Matérn ( ν=23 )
0.40
0.99
0.77
0.97
1.46
Matérn ( ν=25 )
0.07
0.83
0.58
0.8
1.33
TABLE II : Maximum length scales for selected kernels.
Method
Fit ( V-A )
Add ( V-B )
Remove ( V-C )
Predict ( V-A )
N.-W. [ 17 ]
0.12±0.00
-/-
-/-
0.83±0.72
KMP [ 8 ]
12.93±3.93
13.63±3.97
11.42±3.16
1.68±0.33
iKMP [ 20 ]
12.17±3.01
1.18±0.09
0.84±2.33
1.71±0.58
cKMP ( V )
2.66±0.78
0.82±0.16
0.60±0.58
1.35±0.59
TABLE III : Computation times (milliseconds) for adding and removing 3 points and predicting 200 points for letter writing.
Method
Fit ( V-A )
Add ( V-B )
Remove ( V-C )
Predict ( V-A )
N.-W. [ 17 ]
0.22±0.01
-/-
-/-
24.06±2.89
R. KMP ( III )
37.11±3.74
37.20±1.72
34.51±2.61
13.28±1.22
iKMP [ 20 ]
36.22±3.35
4.11±1.60
2.29±0.52
13.15±0.95
cKMP ( V )
7.89±2.32
2.64±0.32
1.27±0.37
8.57±0.82
TABLE IV : Computation times (milliseconds) fitting, adding 3 poses and subsequently removing them in a trajectory with 100 reference poses, predicting 500 poses.
Real-world dynamics shifts pose a critical challenge for reinforcement learning in robotics, as policies tightly coupled to nominal environments often fail catastrophically when physical conditions change. Most existing methods rely on encoding explicitly identified physical parameters into a latent context, a parameter-centric paradigm that depends on pre-specified axes of variation and becomes brittle under unmodeled or compound dynamics changes. We revisit dynamics adaptation from an outcome-centric perspective: rather than telling policies what the dynamics are, we enable them to learn how dynamics affect interaction outcomes. Theoretically, this is grounded in a monotonic relationship between target-domain regret and the Lipschitz constant of a trajectory dynamics encoder. Practically, this constant can be upper-bounded through contrastive learning, yielding a smooth, task-relevant latent topology without privileged dynamics information. On MuJoCo benchmarks, our method consistently outperforms parameter-centric baselines under severe dynamics shifts, including unmodeled and time-varying parameters, while also improving in-distribution stability and latent interpretability. Overall, these results validate that controlling latent geometry is a principled mechanism for robust adaptation.
Zhiming Xu, Weitao Zhou, Xianghui Pan +4
School of Computer Science and Engineering, Tongji University, Shanghai, China · School of Vehicle and Mobility, Tsinghua University, Beijing, China · Simple AI, Beijing, China +1
Applying an imitation learning policy to a new manipulation task usually requires collecting new demonstrations and retraining the model, which makes sample efficiency a practical concern. Pretraining on large-scale robot datasets is effective in this respect, but such datasets are costly to collect and train on, while data augmentation techniques typically require a new round of data generation and retraining for each task. A complementary question is what useful prior can be provided to a policy at negligible cost before any task-specific data are collected. In this study, we construct a geometric visual pretraining dataset in which each scene contains only a plane, an object, and a hand, and trajectories are generated automatically. The scenes contain neither textures nor backgrounds; pretraining primarily exposes the policy to the geometric relationship between the hand and the object. Furthermore, representing the hand as a cube avoids tailoring the dataset to a specific robot morphology. We evaluate this geometric prior using ACT on three simulated robots across five manipulation tasks each, as well as on three real-world robot tasks. Across many of these robot--task combinations, fine-tuning from the geometric prior achieves higher success rates in the early stages of training than training from scratch while using only a small number of task demonstrations. These results suggest that even highly simplified geometric scenes can provide a useful initialization that transfers across robots and to real-world tasks when task data are limited.
Shogo Iwakata, Tomohiro Motoda, Ryosuke Yamada +8
Waseda University · Embodied AI Research Team, AIRC and National Institute of Advanced Industrial Science and Technology (AIST) · Computer Vision Research Team, AIRC and National Institute of Advanced Industrial Science and Technology (AIST) +3
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degrading their broad capabilities. We introduce Proxy Policy Steering (PPS), an inference-time adaptation method that resolves this challenge by training two lightweight proxy policies whose calibrated velocity-space difference steers the frozen base sampler. A reference proxy models the frozen base's behavior on target-task observations, and a task proxy, initialized from the reference, captures how this behavior changes under task supervision. Their difference forms a calibrated velocity-space residual that steers the frozen base sampler at every denoising step. We identify the conditions under which this residual isolates the change induced by task supervision, and validate them empirically. Because the base is never directly modified, its broad capabilities remain available at inference, including behaviors such as recovery from failure that the demonstrations themselves do not exercise. Adaptation requires only forward velocity predictions from the base, making PPS lightweight to train and applicable even without access to the base's parameters. On 8 real-world and 4 simulation manipulation tasks, PPS lifts the state-of-the-art pi 0.5 base policy by 53% absolute success rate on average, with zero-to-one gains on tasks the base never solves, while preserving the base's broad capabilities. PPS outperforms LoRA fine-tuning, from-scratch specialists, residual policies, and prior inference-time steering methods.