PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning
Organizations: NVIDIA
Abstract
Real-world reinforcement learning (RL) systems still struggle with the demands of contact-rich industrial manipulation, including micrometer-level precision, success rates above 99%, and human-level cycle times. Although off-policy algorithms can improve performance by leveraging demonstrations and interventions, a key bottleneck is the lack of an intuitive interface for collecting such guidance while complying with constraints of the physical system and the policy. We propose PAKT, a framework for kinesthetic teaching in RL. As opposed to teleoperation approaches, PAKT relies on kinesthetic guidance, which is widely used in industry. However, a critical weakness of kinesthetic guidance is the possibility for the operator to move the robot along trajectories (e.g., velocities, accelerations, jerk) that the robot and/or policy cannot physically reproduce. Using PAKT, operators guide the robot through admittance control, which maps human-applied forces to motion. The downstream reference generator applies the same kinematic limits used during policy execution, keeping the collected trajectories within these limits. To support this teaching interface with an appropriate execution layer, PAKT adds a high-performance control stack that maps low-frequency RL actions to high-frequency torque commands. It consists of a reference generator and subsequent impedance controller, where the reference generator preserves the tracking performance of the impedance controller while improving contact handling and producing smoother policy actions. Across the reported runs on four insertion and industrial assembly benchmarks, including a data center compute tray, the end-to-end system reduces cycle time by 23%-48% and cumulative intervention count by 62%-86% relative to the HIL-SERL baseline. Project website: https://pakt-website.github.io/pakt-website}{https://pakt-website.github.io/pakt-website
Figures & tables
| ID | Equation | Comment |
|---|---|---|
| C1 | Eq. 10 | Ours |
| C2 | No inertia compensation or desired velocity | |
| C3 | No inertia compensation or desired velocity, and no damping design | |
| C4 | No inertia compensation or desired velocity, and no damping design, but with integrator | |
| C5 | No inertia compensation or desired velocity, and no damping design, but with error clipping |
| Controller | Average tracking error [mm] | Contact force [N] |
|---|---|---|
| C1 (PAKT) | ||
| C2 | ||
| C3 | ||
| C4 | ||
| C5 (HIL-SERL) |
| Task | Initial demonstration time [s] | Success rate | Average cycle time [s] | Cumulative interventions | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| B | H | P | B | H | P | B | H | P | B | H | P | |
| Peg | ||||||||||||
| RAM insertion | ||||||||||||
| Limit fixture | ||||||||||||
| Busbar | ||||||||||||
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| Robot platform | Franka Emika Robot with Franka Hand |
| Control compute | NVIDIA Jetson Orin AGX with L4T and PREEMPT_RT kernel |
| Policy / classifier compute | Desktop PC with Ubuntu 22.04 and NVIDIA RTX 4090 |
| Policy command rate | Hz |
| Cameras | Two Intel RealSense D405 wrist cameras |
| Observation space | Two RGB wrist-camera images, Cartesian pose (zyx Euler convention), Cartesian velocity |
| Task | Episode length [steps] | Action scale (translation [m], rotation [rad]) | reset range [m] | Yaw reset range [rad] |
|---|---|---|---|---|
| Peg insertion | ||||
| RAM insertion | ||||
| Limit fixture | ||||
| Busbar |
| Parameter | Value |
|---|---|
| Learning algorithm | RLPD with SAC-style pixel agent |
| Batch size | total = demo + online RL |
| Critic-to-actor ratio | |
| Discount factor | |
| Maximum training steps | gradient steps |
| Replay buffer capacity | transitions |
| Parameter | Value |
|---|---|
| Policy hidden dimensions | |
| Policy activation | tanh |
| Policy output distribution | Gaussian with squashing |
| Standard-deviation parameterization | exp |
| Standard-deviation range | |
| Policy layer normalization | Enabled |
| Parameter | Value |
|---|---|
| Encoder type | Pretrained frozen ResNet-10 |
| Image pretraining | ImageNet |
| Pre-pooling | Enabled |
| Pooling method | Spatial learned embeddings |
| Number of spatial blocks | |
| Bottleneck dimension |
| Parameter | Value |
|---|---|
| Classifier inputs | Task-relevant wrist-camera images |
| Backbone | Pretrained frozen ResNet-10 |
| Pooling | Spatial learned embeddings |
| Bottleneck dimension | |
| Classification head | Two-layer MLP |
| Loss | Cross-entropy |
| Parameter | Baseline | Ours |
|---|---|---|
| Cartesian translational stiffness | N/m | N/m |
| Cartesian rotational stiffness | N m/rad | N m/rad |
| Cartesian translational damping | N s/m | N/A |
| Cartesian rotational damping | N m s/rad | N/A |
| Cartesian translational damping factors | N/A | |
| Cartesian rotational damping factors | N/A |
| Task | Trans. vel. | Trans. acc. | Rot. vel. | Rot. acc. | Trans. jerk | Rot. jerk |
|---|---|---|---|---|---|---|
| [m/s] | [m/s 2 ] | [rad/s] | [rad/s 2 ] | [m/s 3 ] | [rad/s 3 ] | |
| Peg insertion | ||||||
| RAM insertion | ||||||
| Limit fixture | ||||||
| Busbar |