FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion
Authors: Tao Dong, Jia Yu, Yuxuan Fan, Linna Zhao, Jiaqi Gong, Andong Yang, Chao Gao, Guyue Zhou
Organizations: Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China. · University of Science and Technology Beijing, Beijing, China. · Nanyang Technological University, Singapore. · Department of Electrical Engineering. Tsinghua University, Beijing, China.
Humanoid locomotion over complex terrain requires anticipating footholds that may no longer be visible at touchdown. Limited camera coverage and self-occlusion make it necessary to retrieve relevant terrain information from earlier observations. We present FootQuery, a perceptive locomotion framework that queries depth history using each foot's predicted next touchdown. The policy predicts touchdown locations and uncertainty from proprioception and uses these distributions, together with per-foot features, to query sparsely sampled historical depth frames. During training, realized contacts are projected into historical images to supervise retrieval at the regions where those contacts were visible. The retrieved per-foot features are fused with global visual memory to generate control actions. A progressive force-assistance curriculum supports early exploration, while event-consistent tread-midline shaping encourages coordinated stair contacts. Deployment requires only proprioception and onboard depth images. In simulation, the complete framework outperforms its component ablations on the most challenging tested stairs, gaps, and platforms. Real-world experiments on a Unitree G1 demonstrate continuous traversal with a single policy across outdoor stairs and indoor routes combining stair ascent and descent, platforms, and gaps. These results support organizing visual history around anticipated contacts for perceptive humanoid locomotion.
Figures & tables
Fig. 2: FootQuery and its training support. (a) A future contact region leaves the current view. (b) Predicted per-foot touchdown distributions guide depth-history queries; realized contacts supervise retrieval during training. (c) Progressive force assistance and tread-midline rewards support exploration and successive stair contacts. Distributions and read regions are schematic.
Fig. 3: Historical visibility of a future contact. Left: cyan points show visible terrain during stair ascent; orange crosses mark the realized future right-foot touchdown. Middle/right: depth frames aged 0.02/0.42 s. The touchdown is outside the latest ROI but visible in the older frame. Contact markers are offline references.
Fig. 4: FootQuery architecture. Historical depth supplies global memory and local tokens. Proprioceptive foot features and detached touchdown distributions condition left/right queries. State estimation uses global visual and proprioceptive features; the action head additionally receives foot contexts and detached predictions. Purple bars mark stop-gradient.
Fig. 5: Pelvis assistance during stair ascent. Blue: upward support Fz . Orange: restoring torque τ . The translucent pose shows tilt θ relative to world vertical.
Fig. 6: Tread-midline geometry. (a) Stair contact side view. (b) Selected tread from above: orange edges, green midline, and magenta projected sole center. Travel-direction error is e=∣sproj−c∣=∣uteval−c∣ .
Fig. 7: Success rates for Ours, HPL, and three component ablations across terrain difficulty. Stair results are not separated into ascent and descent. All panels use the same percentage scale.
Fig. 8: All-terrain training success with and without progressive force assistance. One pair of aggregate traces shows success versus PPO iteration.
Fig. 9: Selected stair contacts without (left) and with (right) tread-midline shaping during ascent (top) and descent (bottom). Red dashed boxes identify the evaluated contacts.
Fig. 10: Touchdown predictions on (a) flat ground, (b,c) gaps, (d,e) platforms, and (f,g) stairs. Top: scene views. Bottom: current base-yaw XY coordinates, with +x left and +y down. Cyan circles: predictions; red crosses: realized future contacts recorded offline; gray triangles: base origin. L/R identify the foot.
Fig. 11: Real-world traversal with one policy. (a) Outdoor stair ascent. (b) An indoor route combining stair ascent/descent and platforms. (c) An indoor route combining platforms and gaps. Overlaid robot poses in (b,c) show successive execution stages.
Method
Flat
Rough
Slope
Platform
Gaps
Stairs
All
0–10 cm
20∘
40 cm
45 cm
15 × 30 cm
MoRE
100
100
100
0
100
100
87.5
Hiking
100
100
100
0
0
100
75.0
Ours
100
100
100
100
100
100
100
TABLE I: Success rates (%) in simulation, 10 trials per terrain.
Diagnostic
Value
Current-frame ROI visible
7.82%
History-only ROI visible
60.07%
Any retained ROI frame visible
67.89%
All-head mean positive-set mass
34.30%
Uniform positive-set mass reference
5.58%
GT-best-head Top-1 hit
81.73%
TABLE II: Future touchdown visibility and attention alignment during stair traversal.
While recent advances in perceptive locomotion have enabled humanoid robots to traverse structured terrains, agile parkour in highly discontinuous environments remains an open challenge. In particular, crossing sparse footholds and narrow support regions requires precise foothold selection, effective use of visual observations, and consistent alternating foot placement during fast transitions. In this paper, we present a perceptive humanoid parkour framework that enables stable traversal across terrains with limited foothold availability using only onboard depth observations. The framework features a saliency-guided temporal perception module that combines a saliency prior with gated memory. It retains informative depth features across frames, enabling reliable foot placement from partial observations. By introducing an alternation loss, our symmetry regularization encourages alternating gait patterns and improves traversal robustness. Extensive experiments show that our method significantly improves success rate and foothold accuracy on challenging terrains in both simulation and the real world.
Enabling humanoid robots to operate in complex, dynamic environments remains a critical challenge, fundamentally limited by the ability to navigate robustly, safely, and accurately. While reinforcement learning with velocity-commanded policies has achieved remarkable robustness in humanoid locomotion, this approach lacks explicit control of the foothold placement, leading to unsafe behavior, such as stepping onto human feet, or imprecise navigation, hindering the following manipulation task. Conversely, explicit foothold-tracking policies offer a promising alternative by directly being commanded with target foot poses. However, existing approaches are often limited by unrealistic state assumptions, compromising real-world deployment, or they are part of staged pipelines, making them tied to specific downstream tasks. In this work, we introduce a novel, lightweight framework for training general-purpose 3D foothold-tracking policies. By dynamically providing footstep support through a goal sampler, this method enables the learned policy to be agnostic to specific terrains. Our new target representation effectively mitigates challenges arising in the real world, such as noisy and inaccurate pose estimation and foot contact estimation. Designed for direct real-world transfer, our policy acts as a standalone low-level controller that can be seamlessly paired with various high-level foothold generators. We demonstrate the effectiveness of our framework through extensive experiments in simulation and in the real world. By coupling our policy with different upstream planners, we achieve natural and accurate locomotion in challenging settings, paving the way for loco-manipulation tasks in complex environments.
Alessandro Montenegro, Shihao Li, Puze Liu +2
Politecnico di Milano · Tongji University · Technische Universität Darmstadt +2
Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a key domain gap between humans and humanoid robots: the absence of rich tactile sensing in most humanoid systems. We address this problem with TactileStep, a deployable tactile learning framework that brings sole pressure sensing into humanoid locomotion control for softer touchdowns and more stable support. TactileStep aligns tactile simulation with the real pressure insole, allowing the policy to learn from the same contact features available on hardware. During training, we use tactile and motion cues to recognize different foot-contact phases and apply phase-aware rewards that encourage safer landing and more stable stance. Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.