cs.CVOct 4, 2026

Deep Prior Learning for Embodied Perception

Authors: Yimou Wu, Jiaxin Guo, Yun-hui Liu, Zheng Li

Organizations: The Chinese University of Hong Kong

Abstract

Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

    Apr 20, 2026Kangan Qian, ChuChu Xie, Yang Zhong +13EmbodiedRecent Vision-Language Models

  2. RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

    Jul 31, 2026Qian Wang, Longrui Chen, Peiran Sun +8Visual Geometry Grounded TransformerRobotic Perception

  3. QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer

    May 29, 2026Zhizhen Pan, Hesong Wang, Huan WangVisual Geometry Grounded Transformer3D Perception