TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion
Organizations: Tsinghua University
Abstract
Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.
Figures & tables
| Metric | Flat | Slope | Upstairs | Downstairs | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TactileStep | w/o stable | Baseline | TactileStep | w/o stable | Baseline | TactileStep | w/o stable | Baseline | TactileStep | w/o stable | Baseline | |
| Contact Area Ratio | 0.929 | 0.805 | 0.536 | 0.545 | ||||||||
| CoP Margin [mm] | 35.79 | 34.24 | 29.05 | 29.43 | ||||||||
| Terrain | Success Rate | Velocity RMSE [m/s] | Traversal Time [s] | Energy [J] | ||||
|---|---|---|---|---|---|---|---|---|
| TactileStep | Baseline | TactileStep | Baseline | TactileStep | Baseline | TactileStep | Baseline | |
| Stair up | 100% | 99.98% | ||||||
| Stair down | 100% | 99.98% | ||||||
| Platform up | 100% | 99.93% | ||||||
| Platform down | 100% | 92.26% | ||||||
| Flat | 99.98% | 99.19% | ||||||
| Terrain | [N] | [dB] | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Baseline | w/o tac. obs. | TactileStep | Baseline | w/o tac. obs. | TactileStep | Baseline | w/o tac. obs. | TactileStep | |
| Stair up | 349.8 | 72.2 | 0.482 | ||||||
| Stair down | 344.1 | 67.1 | 0.598 | ||||||
| Platform up | 355.7 | 83.0 | – | – | – | ||||
| Platform down | 404.6 | 75.9 | – | – | – | ||||
| Flat | 191.3 | 66.5 | 0.510 | ||||||
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Terrain | [N] | [mm] | ||||
|---|---|---|---|---|---|---|
| Dual critic | Single critic | Dual critic | Single critic | Dual critic | Single critic | |
| Flat | 278.16 | 307.23 | 35.79 | 34.33 | 0.929 | 0.824 |
| Slope | 297.32 | 335.90 | 34.24 | 31.40 | 0.805 | 0.635 |
| Stair descent | 320.74 | 579.83 | 29.43 | 28.08 | 0.545 | 0.516 |
| Stair ascent | 304.45 | 402.71 | 29.05 | 28.63 | 0.536 | 0.488 |
| Platform drop-down | 310.62 | 333.93 | – | – | – | – |
| Reward | Weight / scale | Formula | Activation | Purpose |
|---|---|---|---|---|
| Per-step | Track planar velocity | |||
| Per-step | Track yaw rate | |||
| Per-step | Penalize excessive yaw-rate command magnitude | |||
| Per-step | Encourage stable rollouts | |||
| Per-step | Penalize hip yaw/roll deviation | |||
| Per-step | Suppress base roll/pitch angular motion |
| Reward | Weight / scale | Formula | Activation | Purpose |
|---|---|---|---|---|
| Command-gated | Discourage stopping under forward commands | |||
| Target event | Reward reaching the target | |||
| Command-gated | Regulate standing posture | |||
| if exactly one foot is in contact, else | Contact-pattern gated | Encourage alternating swing and stance timing | ||
| Contact-gated | Penalize foot sliding during contact | |||
| Contact-gated | Encourage the sole to stay close to local terrain |
| Terrain category | Task setting | Curriculum parameter | Purpose |
|---|---|---|---|
| Rough flat | Forward locomotion on uneven ground | Perlin height scale: – | Robust walking under mild contact-height perturbations |
| Rough standing | Standing on uneven ground | Perlin height scale: – | Static support stability and CoP regulation |
| Stair descent | Walking down stairs | Step height: – | Soft touchdown and impact reduction during downward transitions |
| Stair ascent | Walking up stairs | Step height: – | Swing clearance and foothold establishment during upward transitions |
| Platform drop-down | Stepping down from high platforms | Platform height: – | High-impact landing control and stance recovery |
| Platform step-up | Stepping onto high platforms | Platform height: – | Large-step clearance, foot placement, and push-off coordination |
| Category | Configuration |
|---|---|
| Environment | environments; rollout steps per environment; transitions per update; training iterations. |
| Policy | Encoder MoE actor-critic with experts; actor MLP ; critic MLP ; ELU activation; initial action std. . |
| Depth encoder | CNN output dimension ; channels ; kernel/stride/padding ; MLP hidden dimensions . |
| PPO | AdamW optimizer; learning rate with adaptive KL schedule; desired KL ; mini-batches; mini-batch size ; epochs per update; clip range ; ; ; entropy coefficient ; value loss coefficient ; max gradient norm . |
| Dual critic | Two value heads estimate dense and sparse environment returns, respectively; advantage mixing weights . |
| AMP | Discriminator MLP with ReLU; AdamW optimizer with learning rate ; discriminator reward coefficient ; gradient penalty coefficient ; weight decay ; logit decay . |
| Category | Randomized quantity | Range / distribution |
| Contact material | Static friction | |
| Dynamic friction | , with | |
| Restitution | ||
| Robot model | Default joint-position offset | |
| Torso CoM offset, | ||
| Torso CoM offset, |
| Noise source | Value | Description |
|---|---|---|
| Relative taxel force error | Per-taxel scaling | |
| Taxel XY offset std. | Episode-level taxel-coordinate perturbation | |
| Taxel XY offset clip | Maximum coordinate perturbation | |
| Delay probability | Probability of using delayed tactile history | |
| Maximum delay | frame | At most one tactile update delay |
| Force clipping | Clamp measured taxel force to be nonnegative |