Safe and efficient quadruped navigation over unfamiliar terrain requires predicting terrain-robot interaction before contact: geometry and visual appearance alone cannot reveal how the robot will slip, load its feet, or expend energy. This paper presents a continual learning pipeline that uses locomotion experience to learn these interaction outcomes from pre-contact images and continually updates the predictions as new contacts are observed. Pre-contact descriptors, produced by a DINOv3 backbone model frozen during training, are mapped to five proprioceptive indicators weighted according to measurement reliability: planar foot slip, mean normal ground-reaction force, traction index, cost of transport, and touchdown loading rate. A compact evidential regressor allows us to predict these indicators together with aleatoric and epistemic uncertainty from the visual descriptors. Continual adaptation combines bounded experience replay with a validation gate: candidate models replace the deployed predictor only when they improve performance on recent held-out data while keeping degradation on historical held-out data within a prescribed tolerance. Predictions and epistemic uncertainty are projected into a local multilayer map and combined into a conservative traversability score map whose property weights can be adjusted without retraining. The resulting map is used for downstream navigation tests. The ROS2 implementation supports evaluation on a Unitree Go2 in simulation and on hardware, with models trained separately in each domain. On a sequential hardware stream over three previously unseen terrains, gated replay reduces final anchor negative log-likelihood (NLL) degradation by 23.1% relative to replay without the gate while attaining similar new-terrain adaptation.
Figures & tables
Fig. 1: Qualitative hardware dataflow: RGB features are extracted by frozen DINOv3, converted by the evidential regressor into five interaction-property score images, and projected into a traversability grid map. The foam surfaces receive lower traversability scores. Artificial grass scores close to the laboratory floor, but slightly lower because it is perceived as more slippery.
Method
Predicted output
Online
Memory / retention
Uncertainty / confidence
Gate
WVN [ 10 ]
Traversability score
Yes
Mission graph
Reconstruction-based confidence
–
Chen et al. [ 3 ]
Friction, stiffness
Yes
Mission graph
Anomaly-based confidence mask
–
EVORA [ 15 ]
Linear/angular traction
–
–
Evidential distributions + feature density (A/E)
–
SALON [ 11 ]
Roughness, speed
Yes
Experience buffer
GP predictive variance
–
IMOST [ 12 ]
Traversability
Yes
Incremental memory
Reconstruction-based anomaly / sampling scores
–
Lee et al. [ 25 ]
Linear/angular traction
Yes
Generative recall
Probabilistic ensemble (A/E)
–
TABLE I: Comparison of selected traversability methods. Online denotes predictor adaptation from newly acquired deployment experience. Gate denotes candidate acceptance based jointly on improved predictive performance on recent held-out data and a prescribed tolerance on historical held-out degradation relative to the deployed model. A/E: aleatoric/epistemic uncertainty; –: not reported as a component of the described method.
Fig. 2: Overview of the proposed experience-driven traversability pipeline. The deployed model performs uninterrupted inference while a separate candidate is trained from newly acquired visual–proprioceptive experience.
Domain
Indicator (units)
MAE
R2
C80
Sim.
Planar foot Slip (m)
0.0050
0.3958
82.11%
Traction index (–)
0.0711
0.6037
74.80%
Real
Planar foot Slip (m)
0.0047
0.4110
80.21%
Normal GRF (N)
3.6191
0.3470
80.21%
Traction index (–)
0.0296
0.5269
82.81%
CoT (–)
0.235608
0.3128
85.42%
TABLE II: Offline prediction on held-out contact-level tests. Within each domain, a separate multi-terrain acquisition dataset was divided into 70% training, 15% validation, and 15% testing. Contacts sharing a timestamp or identical descriptor were grouped. MAE is in the shown units and C80 is a percentage (nominal: 80).
Method
Adapt. ↑
Lab ↓
Anchor ↓
Final MAE ↓
(%)
(%)
NLL (%)
CL-P
1.658
0.801
0.608
1.03711
Replay
1.664
0.924
0.790
1.03633
TABLE III: Sequential hardware results, averaged over five optimization seeds on one fixed stream. Lab uses an independent test; anchor NLL uses the gate-validation set, not an independent test. Final MAE is the four-domain macro standardized MAE.
Fig. 3: Relative anchor NLL of deployed models during artificial grass → textured foam → foam. Curves show five-seed means; bands show seed ranges, crosses gate-rejected CL-P candidates, and vertical lines phase transitions. The anchor is used for candidate selection.
Fig. 4: Traversability-aware navigation in simulation: environment and camera view (left), predicted interaction-property scores (center), and resulting navigation costmap and planned path (right).
Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods often fail to adapt to novel terrains. Even when adaptation is achieved, retaining experience from previously trained environments remains a challenge, a problem known as catastrophic forgetting. To address this challenge, we propose a continual learning framework for traversability prediction that incrementally adapts to new terrains using a generative experience recall model. A key virtue of the proposed framework is two folds: i) retain prior experience without storing past data; and ii) incorporate the uncertainty of the generated samples from the recall model, enabling uncertainty-aware adaptation. Real-world experiments with a skid-steering robot validate the effectiveness of the proposed framework, demonstrating its ability to adapt across a series of diverse environments while mitigating catastrophic forgetting.
Hojin Lee, Yunho Lee, Daniel A Duecker +1
Ulsan National Institute of Science and Technology, Ulsan, 44919, Republic of Korea · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM), Germany
Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective. We instead treat these functions through distinct mechanisms: rewards for task specification, constraints for operational limits, energy minimization for gait preference, and exteroceptive perception for adapting energy use to terrain difficulty. We show that these components jointly enable efficient, terrain-adaptive locomotion, and that removing each component exposes a distinct failure mode. Our formulation removes explicit gait priors (including air-time, contact-count, and foot-clearance targets) in favor of emergent behavior. Compared to a conventional complex-reward baseline, our formulation achieves comparable terrain traversal while reducing cost of transport by 56% and operational-limit violations by 96%. The resulting policies transfer zero-shot to a physical Unitree Go2 using LiDAR-based elevation mapping. Project website with videos: https://tinyurl.com/locomposition.
Loukas Kordos, Leonard T. Franz, Simon Rappenecker +4
Technical University of Munich. · University of Tübingen. · Hertie Institute for Clinical Brain Research & Center for Integrative Neuroscience. +1
Learning-based control has revolutionized dynamic locomotion, yet navigating unstructured terrain remains limited by a robot's incomplete awareness of imminent ground contact. While global perception systems such as LiDARs and depth cameras provide environmental context, they are frequently plagued by latencies, occlusions, and the high computational cost of dense geometric reconstruction. On the other hand, proprioceptive feedback is purely reactive, initiating corrections only after impact has occurred. This work explores embedding a minimal suite of low-cost, high-frequency infrared proximity sensors directly into the feet of a quadrupedal robot. These sensors provide "pre-contact" feedback that is robust to self-occlusions and significantly less computationally demanding than conventional vision-based pipelines. By integrating these localized signals into a reinforcement learning framework, we enable the robot to anticipate terrain discontinuities such as gaps and stepping stones that are problematic for traditional perception stacks due to occlusions or state estimation drift. We demonstrate that such sparse, near-field sensing can be reliably modeled in simulation and transferred to the real world with high fidelity. Experimental results show that local proximity sensing substantially improves traversal robustness over discrete terrain and offers a low-power, low-latency alternative or complement to complex global perception suites in unpredictable environments. For more information about results and methods, please see the project website: https://sites.google.com/view/foot-tof/home.
Jiale Fan, Connor Flynn, Tianao Xu +4
All authors are with the Robotic Systems Lab, ETH Zurich.