Safe and efficient quadruped navigation over unfamiliar terrain requires predicting terrain-robot interaction before contact: geometry and visual appearance alone cannot reveal how the robot will slip, load its feet, or expend energy. This paper presents a continual learning pipeline that uses locomotion experience to learn these interaction outcomes from pre-contact images and continually updates the predictions as new contacts are observed. Pre-contact descriptors, produced by a DINOv3 backbone model frozen during training, are mapped to five proprioceptive indicators weighted according to measurement reliability: planar foot slip, mean normal ground-reaction force, traction index, cost of transport, and touchdown loading rate. A compact evidential regressor allows us to predict these indicators together with aleatoric and epistemic uncertainty from the visual descriptors. Continual adaptation combines bounded experience replay with a validation gate: candidate models replace the deployed predictor only when they improve performance on recent held-out data while keeping degradation on historical held-out data within a prescribed tolerance. Predictions and epistemic uncertainty are projected into a local multilayer map and combined into a conservative traversability score map whose property weights can be adjusted without retraining. The resulting map is used for downstream navigation tests. The ROS2 implementation supports evaluation on a Unitree Go2 in simulation and on hardware, with models trained separately in each domain. On a sequential hardware stream over three previously unseen terrains, gated replay reduces final anchor negative log-likelihood (NLL) degradation by 23.1% relative to replay without the gate while attaining similar new-terrain adaptation.
Figures & tables
Fig. 1: Qualitative hardware dataflow: RGB features are extracted by frozen DINOv3, converted by the evidential regressor into five interaction-property score images, and projected into a traversability grid map. The foam surfaces receive lower traversability scores. Artificial grass scores close to the laboratory floor, but slightly lower because it is perceived as more slippery.
Method
Predicted output
Online
Memory / retention
Uncertainty / confidence
Gate
WVN [ 10 ]
Traversability score
Yes
Mission graph
Reconstruction-based confidence
–
Chen et al. [ 3 ]
Friction, stiffness
Yes
Mission graph
Anomaly-based confidence mask
–
EVORA [ 15 ]
Linear/angular traction
–
–
Evidential distributions + feature density (A/E)
–
SALON [ 11 ]
Roughness, speed
Yes
Experience buffer
GP predictive variance
–
IMOST [ 12 ]
Traversability
Yes
Incremental memory
Reconstruction-based anomaly / sampling scores
–
Lee et al. [ 25 ]
Linear/angular traction
Yes
Generative recall
Probabilistic ensemble (A/E)
–
TABLE I: Comparison of selected traversability methods. Online denotes predictor adaptation from newly acquired deployment experience. Gate denotes candidate acceptance based jointly on improved predictive performance on recent held-out data and a prescribed tolerance on historical held-out degradation relative to the deployed model. A/E: aleatoric/epistemic uncertainty; –: not reported as a component of the described method.
Fig. 2: Overview of the proposed experience-driven traversability pipeline. The deployed model performs uninterrupted inference while a separate candidate is trained from newly acquired visual–proprioceptive experience.
Domain
Indicator (units)
MAE
R2
C80
Sim.
Planar foot Slip (m)
0.0050
0.3958
82.11%
Traction index (–)
0.0711
0.6037
74.80%
Real
Planar foot Slip (m)
0.0047
0.4110
80.21%
Normal GRF (N)
3.6191
0.3470
80.21%
Traction index (–)
0.0296
0.5269
82.81%
CoT (–)
0.235608
0.3128
85.42%
TABLE II: Offline prediction on held-out contact-level tests. Within each domain, a separate multi-terrain acquisition dataset was divided into 70% training, 15% validation, and 15% testing. Contacts sharing a timestamp or identical descriptor were grouped. MAE is in the shown units and C80 is a percentage (nominal: 80).
Method
Adapt. ↑
Lab ↓
Anchor ↓
Final MAE ↓
(%)
(%)
NLL (%)
CL-P
1.658
0.801
0.608
1.03711
Replay
1.664
0.924
0.790
1.03633
TABLE III: Sequential hardware results, averaged over five optimization seeds on one fixed stream. Lab uses an independent test; anchor NLL uses the gate-validation set, not an independent test. Final MAE is the four-domain macro standardized MAE.
Fig. 3: Relative anchor NLL of deployed models during artificial grass → textured foam → foam. Curves show five-seed means; bands show seed ranges, crosses gate-rejected CL-P candidates, and vertical lines phase transitions. The anchor is used for candidate selection.
Fig. 4: Traversability-aware navigation in simulation: environment and camera view (left), predicted interaction-property scores (center), and resulting navigation costmap and planned path (right).
Ulsan National Institute of Science and Technology, Ulsan, 44919, Republic of Korea · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM), Germany