Navigation in vision-denied environments is challenging for humanoid robots because proprioceptive odometry drifts and localization uncertainty accumulates rapidly. We present TAPNAV, a tactile active-perception framework that enables humanoid navigation toward a goal by actively probing surrounding structures without relying on vision. TAPNAV maintains a pose belief from odometry, IMU, and tactile contact observations, and couples uncertainty-aware global route planning with information-gain-driven local probing. The global planner searches for routes that keep predicted localization uncertainty bounded by exploiting opportunities for tactile correction, while the local planner selects probe actions that maximize expected information gain. A whole-body controller coordinates the humanoid's locomotion and end-effector contact to execute the planned navigation and probe motions. We evaluate TAPNAV in simulation and on a Unitree G1 across different floor plans and obstacle geometries. TAPNAV achieves lower state estimation error and a higher task completion rate than baselines. These results demonstrate that actively planning physical interactions with the environment can provide localization cues for reliable humanoid navigation without vision.
Figures & tables
Fig. 2: Overview of the TapNav framework. Odometry, IMU, and tactile observations together update the robot’s SE(2) pose belief. The global planner computes an uncertainty-bounded route with predicted tactile corrections, the local planner selects informative probes based on expected information gain, and the low-level controller executes locomotion and probing.
Fig. 3: Tactile end-effector design and contact-processing pipeline. The tactile image is aggregated into a total activation T(t) , which is calibrated to a normal force and converted to a binary contact outcome using a force threshold.
Fig. 4: Tactile likelihood. (a) The residual δt(i) under hypothesis i and the nearest point on the map surface cta defined in Sec. IV-C ; (b) Under hypothesis 1 , the end-effector would penetrate the surface ( δt(1)<0 ); (c) Under hypothesis 2 , it would not reach the surface ( δt(2)>0 ).
Fig. 5: Information gain of one probe action. (a) The margin uncertainty, modeled as m~ta∼N(μmta,σmta2) . (b) Pose belief along nta before the probe, after a contact, and after a miss (prior truncated at (nta)⊤μtxy−μmta ).
Fig. 6: The low-level controller is composed of an upper- and a lower-body policy. The upper-body policy tracks the target end-effector pose and force, while the lower-body policy tracks footstep commands. The two policies are jointly trained with shared whole-body proprioception.
Fig. 7: TapNav navigation performance in simulation. (a)–(d) correspond to four evaluated maps: (a) a curved corridor, (b) a partitioned room, (c) a zigzag corridor, and (d) a room with scattered obstacles. Each map has its own start (purple) and goal (yellow star). For each map, (i) shows the map layout with the estimated trajectory (blue), the ground-truth trajectory measured from the simulator (black), and contact points from probing (red), where each grid cell is 1m×1m ; (ii), (iii) show position and yaw uncertainty (blue) and error (black) over time, respectively, with vertical red lines marking contact events. The red and yellow boxes highlight the largest decreases in both estimation error and uncertainty, which occur when the robot contacts non-planar wall surfaces.
Parameter
Setting
Controller Frequency
200 / 50 Hz
Replanning Deviation Margin
>0.1 m
Global Planner Grid Size
0.1 m
Minimal Wall Clearance
0.2 m
∣Δxt∣max
0.40 m
(σˉp,σˉψ)
(0.18 m, 20∘ )
TABLE I: Simulation and navigation parameters.
Method
Map a
Map b
Map c
Map d
PosErr
YawErr
SR
PosErr
YawErr
SR
PosErr
YawErr
SR
PosErr
YawErr
SR
(cm) ↓
(deg) ↓
(%) ↑
(cm) ↓
(deg) ↓
(%) ↑
(cm) ↓
(deg) ↓
(%) ↑
(cm) ↓
(deg) ↓
(%) ↑
TapNav
6.5±2.5
1.4±0.4
95
9.4±3.6
2.8±1.8
100
5.9±1.9
2.0±1.3
100
16.8±5.0
2.3±1.6
85
Random-touch
9.6±3.8
2.1±1.4
90
11.4±3.9
2.9±1.4
70
8.7±3.6
2.9±1.7
75
16.6±6.0
2.7±1.8
75
Sweep-touch
11.7±4.5
2.0±1.1
50
11.2±3.9
2.5±1.5
50
10.9±4.2
2.4±1.6
30
13.7±3.5
1.32±0.1
25
Odometry-only
23.4±11.2
3.5±2.9
35
36.5
4.0
5
13.2
0.6
5
28.5±12.2
4.0±4.6
45
TABLE II: MuJoCo simulation results. The navigation performance of TapNav and the baselines is evaluated across four maps. Errors are reported as mean ±std ; entries without a standard deviation correspond to settings in which only a single trial succeeded.
Fig. 8: Yaw-noise sensitivity. (a) position error; (b) yaw error; (c) success rate. Bands and bars show mean ± standard error in (a)–(b) and 95% confidence intervals in (c).
Fig. 9: Simulated open-loop goal tracking with velocity- and footstep-tracking policies (30 trials per policy). (a) Executed trajectories, highlighting the trial with median final position error for each policy; (b) Actual and commanded heading over time.
Fig. 10: TapNav deployed on a Unitree G1 humanoid equipped with our tactile end-effectors, demonstrating real-world navigation capabilities in (a) a kitchen environment and (b) an office hallway. For each environment, (i) shows snapshots of the real-world deployment; (ii) shows the robot’s path and active tactile probes in the robot’s egocentric view.
Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision and struggle with autonomous tactile exploration and object identification. We present TACTFUL, a vision-free tactile exploration framework that enables a multi-fingered robot to autonomously explore confined workspaces, discover objects through contact, and identify them via tactile reconstruction. Trained entirely on real hardware without simulation, our system learns a single policy that balances global workspace exploration with local surface refinement through a dynamic reward schedule. Our results demonstrate that tactile sensing, when paired with structured learning, can serve as an effective primary modality for object-level reasoning, achieving 77% success with 0.015 m average reconstruction error and outperforming baseline approaches on real-world objects.
Shivani Kamtikar, Chung Hee Kim, Camilla Tabasso +3
Amazon Fulfillment Technologies & Robotics, Westborough, MA, USA. · Siebel School of Computing and Data Science, University of Illinois at Urbana-Champaign, Champaign, IL, USA. · Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, USA.
Active perception allows autonomous agents to select their viewpoints rather than passively process the viewpoints given to them, enabling them to target where to reduce uncertainty about their environment. Learned systems typically encourage this behavior with hand-designed proxy objectives, such as coverage or curiosity bonuses, that may conflict with the task. In this work, we propose a method to learn emergent active perception (LEAP) without augmentation of the task objective. We formulate the problem of goal-oriented navigation over hazardous terrains with goals that must be discovered visually. We then propose an architecture for navigation policies with active perception, and train them on a terrain curriculum where task pressure alone leads to the emergence of gaze control. Key to this emergence, LEAP works on a gaze-invariant representation that integrates depth images into egocentric belief maps. We validate its performance in held-out evaluation scenarios, where it achieves a 92.7% success rate, compared to 74.2% for scripted or 34.5% for passive perception, and comes within 4.6 points of a privileged oracle. We validate that LEAP navigation policies, unchanged, can be directly applied to steering quadrupedal locomotion policies in physics simulation.
Ü. Bora Gökbakan, Stéphane Caron, Philippe Souères
WILLOW · Inria and DI ENS, PSL, Paris, France · ISIR +3
Agile humanoid locomotion across diverse challenging terrain demands both wide perceptual coverage and precise local geometry understanding. Motivated by the way humans selectively look at relevant terrain during locomotion, we introduce TAGA, a Terrain-aware Active Gaze learning framework for Attention-based humanoid control. By fusing vision, proprioception, and motion commands, our framework guides the model to learn anticipatory cues and actively attend to specific areas of the height scan, selectively using these informative regions for the downstream network. This adaptively increases the information density of observations under tight onboard computational constraints, thus enabling fine-grained perceptive locomotion over larger-scale terrains. We find that such gaze behaviors can naturally emerge through reinforcement learning alone, without requiring additional supervision or explicit guidance, significantly improve training efficiency. As a result, the trained policy demonstrates robust and generalizable locomotion in simulation and on hardware, including reliable terrain-aware foothold selection, elevated-platform traversal, competitive sparse-foothold traversal, and the largest reported real-world gap traversal distance of 1.2m among perceptive humanoid locomotion systems, while maintaining stability under severe perceptual disturbances and environmental interference.
Peizhuo Li, Hongyi Li, Mingfeng Fan +9
MarmotLab, National University of Singapore · Center of X-Mechanics, Zhejiang University · South China University of Technology