A wide range of approaches have been developed for camera pose estimation, including correspondence-based methods, end-to-end pose regression, and recent 3D geometric foundation models. Our key observation is that no single estimator is optimal for diverse challenges, such as wide baselines, lack of texture, appearance changes, and occlusions. Further analysis reveals substantial performance variation across both benchmarks and individual image pairs, with different estimators exhibiting complementary strengths. We introduce PoseAgent, an agentic framework for relative camera pose estimation that dynamically orchestrates pose estimators through learnable ranking and verification. Given an image pair, a profiling agent first extracts appearance, semantic, and geometric features relevant to pose estimation, e.g., scene type. A learned ranking agent then predicts the relative competence of multiple pose estimators given the image-pair profile. The top-ranked estimator is executed, and its predicted pose is assessed by a learned verification agent that estimates the corresponding pose error. When verification fails, PoseAgent adaptively invokes lower-ranked estimators until a candidate is accepted or the execution budget is reached. For pose verification, our verification network predicts pose errors more accurately than prior models. For pose estimation, PoseAgent improves AUC@5 degree up to 4.2% over the strongest standalone estimator on each of ARKitScenes, MegaDepth, ScanNet++, and RealEstate10K. On ARKitScenes, PoseAgent also outperforms VLM-based agents, which include a VLM ranker with the same verifier and fallback policy. These results demonstrate the effectiveness of our learned ranking and verification.
Figures & tables
Figure 1: We propose PoseAgent , an agentic framework that adaptively ranks and executes pose estimators and verifies pose estimation with orchestration. Our framework is a meta-learner and improves a set of standalone pose estimators.
Figure 2: Distributions of best estimator across various datasets. No single estimator dominates the lowest pose error across all datasets or image pairs.
Figure 3: PoseAgent profiles an image pair, ranks and sequentially executes candidate estimators, and verifies each prediction until a pose is accepted following a fallback policy.
Figure 4: Architecture of Pose Verification Network (PVN). The network is conditioned on camera intrinsics and a candidate pose to predict rotation and translation errors.
ARKitScenes
MegaDepth
ScanNet++
RealEstate10K
Method
@5°
@10°
@20°
@5°
@10°
@20°
@5°
@10°
@20°
@5°
@10°
@20°
Reloc3r
0.456
0.675
0.819
0.562
0.732
0.847
0.684
0.833
0.915
0.605
0.764
0.861
VGGT
0.243
0.455
0.647
0.530
0.687
0.804
0.194
0.342
0.515
0.323
0.539
0.712
SG-I
0.161
0.296
0.451
0.162
0.257
0.355
0.094
0.184
0.299
0.359
0.524
0.662
SG-O
0.242
0.394
0.539
0.430
0.572
0.686
0.258
0.378
0.485
0.568
0.706
0.802
SG-I-RS
0.146
0.264
0.393
0.320
0.413
0.496
0.152
0.227
0.308
0.467
0.596
0.696
Table 1: Relative pose AUC on four datasets at 5°/10°/20° . Bold indicates the best result, and underline indicates the best standalone estimator.
Method
ARKitScenes
MegaDepth
ScanNet++
RealEstate10K
Rank-1 (Ranker-Only)
0.460 / 0.674 / 0.816
0.682 / 0.794 / 0.869
0.702 / 0.839 / 0.916
0.729 / 0.836 / 0.901
Rank-2
0.393 / 0.591 / 0.742
0.655 / 0.774 / 0.855
0.543 / 0.680 / 0.779
0.682 / 0.805 / 0.882
Verifier-Only
0.455 / 0.669 / 0.813
0.668 / 0.794 / 0.879
0.728 / 0.856 / 0.927
0.698 / 0.820 / 0.894
Fixed-order Ranking + PVN
0.467 / 0.681 / 0.821
0.611 / 0.762 / 0.863
0.694 / 0.838 / 0.917
0.636 / 0.786 / 0.875
PoseAgent (Our Ranker + PVN)
0.473 / 0.685 / 0.823
0.693 / 0.810 / 0.889
0.726 / 0.855 / 0.926
0.736 / 0.843 / 0.907
Oracle@2
0.532 / 0.727 / 0.851
0.715 / 0.819 / 0.888
0.762 / 0.876 / 0.937
0.774 / 0.867 / 0.921
Table 2: Ranking and model selection. Results are AUC@5°/10°/20°. Verifier-Only runs all 9 estimators and picks the lowest predicted error. Fixed-Order + PVN uses one estimator order for all datasets (average win rate from train set) with the same PVN, fallback policy, and K=4 .
Dataset
Method
mAA@ 10∘↑
Median error ( ∘ ) ↓
MAE ( ∘ ) ↓
eR/et/emax
eR/et/emax
eR/et/emax
MegaDepth
FSNet
0.850 / 0.683 / 0.655
0.835 / 1.867 / 2.185
5.463 / 6.615 / 8.588
PVN
0.964 / 0.850 / 0.839
0.393 / 0.814 / 1.017
0.787 / 4.681 / 4.243
ScanNet++
FSNet
0.462 / 0.340 / 0.284
5.311 / 10.005 / 13.978
34.147 / 22.934 / 41.854
PVN
0.958 / 0.928 / 0.904
0.599 / 0.712 / 1.034
0.715 / 4.316 / 2.853
Table 3: Pose error prediction and selection on MegaDepth and ScanNet++.
Method
ScanNet1500
MegaDepth1500
Rank-1
0.359 / 0.581 / 0.748
0.700 / 0.814 / 0.891
Order 1 + PVN
0.366 / 0.592 / 0.760
0.715 / 0.828 / 0.902
Order 2 + PVN
0.370 / 0.596 / 0.764
0.716 / 0.828 / 0.903
Oracle@9
0.522 / 0.716 / 0.843
0.818 / 0.897 / 0.944
Table 4: Verification-guided fallback on benchmarks. Each entry reports AUC at 5°/10°/20° . We evaluate two fixed estimator orders while keeping the PVN, verification thresholds, fallback policy, and execution budget unchanged.
Method
AUC@5°/10°/20°
VLM Coding Agent
0.090 / 0.163 / 0.259
PoseAgent
0.473 / 0.685 / 0.823
Table 5: Results on ARKitScenes.
Table 10
Figure 5: Effect of the execution budget K on relative pose accuracy. We use τR=0.5∘ and τt=6∘ in this experiment. Dashed lines indicate the default setting with K=4 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Datasets
Scene Type
Training Pairs
Test Pairs
ARKitScenes
Indoor
20,000
1,095
ScanNet++
Indoor
18,202
1,798
MegaDepth
Outdoor
17,879
2,121
RealEstate10K
Indoor & Outdoor
20,000
5,449
Appendix
Table 8: Datasets for PoseAgent. It uses four datasets covering indoor and outdoor environments.
Tool
Extracted information
OpenCV
Image quality and texture (Laplacian, FAST); ORB matches and RANSAC geometric inliers; Farneback optical flow variation.
CLIP
Zero-shot scene category, scene type
DINOv2
Appearance change measured from the cosine distance between image embeddings.
SegFormer
Semantic category pixel ratios, scene composition
LoFTR
Spatial coverage of confident matches, visual overlap
SuperPoint
Keypoints and descriptor matching statistics, texture and match quality
Appendix
Table 9: Computer vision tools used to construct image-pair profiles.
Figure 6: Performance across predicted estimator rank and oracle top- k selection. Rank- k reports the AUC@10°using the estimator at predicted rank k ; Oracle@ k selects the best prediction among the top- k ranked estimators.
Figure 7: Successful early exit. PoseAgent ranks Reloc3r first and PVN accepts its estimate. The execution terminates after a single model. This example shows how accurate ranking combined with the verification avoids unnecessary model executions.
Figure 8: Recovery from inaccurate top-ranked candidates. PVN rejects the first two candidates, RoMa (Outdoor) and RoMa (Indoor). It then accepts Rank-3 Reloc3r. PoseAgent reduces the final error from the Rank-1 error of 37.35° to 2.59° after three model executions.
Figure 9: Recovery through fallback. None of the Top-4 candidate models satisfies the strict PVN acceptance thresholds, so PoseAgent executes all four models before invoking fallback. Based on the predicted errors, fallback policy selects the earlier Rank-3 RoMa (Indoor) estimate, reducing the actual error from the Rank-1 error of 31.66° to 1.40° . This example shows how fallback can recover a useful pose even when no candidate is directly accepted.
Figure 10: Failure caused by verifier prediction accuracy. PVN rejects the Rank-3 model’s estimate despite its low actual error. It subsequently accepts Rank-4 Reloc3r. This failure example indicates that PVN underestimates the pose error on good candidates, but overestimate on poor ones.
Dataset
Profiling configuration
AUC@5°
AUC@10°
AUC@20°
ARKitScenes
Scene Features
0.467
0.680
0.821
Geometry Features
0.473
0.685
0.824
All features
0.473
0.685
0.823
ScanNet++
Scene Features
0.694
0.838
0.917
Geometry Features
0.726
0.855
0.926
All features
0.726
0.855
0.926
Appendix
Table 10: Ablation of profiling features . Each ranker is trained using the specified feature subset, while PVN, verification thresholds, the fallback policy, and the maximum execution budget ( K=4 ) are kept unchanged. Best results are shown in bold, including ties at the reported precision.
This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alternative to the current trend of fully supervised training with 3D annotations. Concretely, we employ inverse- and forward-dynamics models to learn latent action representations, similar to Genie from large-scale driving videos. Our idea is simple yet effective. Existing methods use latent actions in their original capacity, that is, as action conditioning of world-models or as proxies of robot action parameters in policy networks. Our method, dubbed LA-Pose, repurposes the latent action features as inputs to a camera pose estimator, finetuned on a limited set of high-quality 3D annotations. This formulation enables accurate and generalizable pose prediction while maintaining feed-forward efficiency. Extensive experiments on driving benchmarks show that LA-Pose achieves competitive and even superior performance to state-of-the-art methods while using orders of magnitude less labeled data. Concretely, on the Waymo and PandaSet benchmarks, LA-Pose achieves over 10% higher pose accuracy than recent feed-forward methods. To our knowledge, this work is the first to demonstrate the power of inverse-dynamics self-supervised learning for pose estimation.
Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In this work, we show that passive scene sounds provide cues complementary to vision for relative camera pose estimation for in-the-wild videos. We introduce a simple but effective audio-visual framework that integrates direction-of-arrival (DOA) spectra and binauralized embeddings into a state-of-the-art vision-only pose estimation model. Our results on two large datasets show consistent gains over strong visual baselines, plus robustness when the visual information is corrupted. To our knowledge, this represents the first work to successfully leverage audio for relative camera pose estimation in real-world videos, and it establishes incidental, everyday audio as an unexpected but promising signal for a classic spatial challenge. Project: http://vision.cs.utexas.edu/projects/av_camera_pose.
Daniel Adebi, Sagnik Majumder, Kristen Grauman
The University of Texas at Austin, Austin TX 78712, USA
Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-aware methods exist, they are often not unified: semantic-based approaches are brittle, per-sequence optimization methods fail on short sequences, and other learned models may degrade on static-only scenes. We present WildPose, a unified monocular pose estimation framework that is robust in dynamic environments while maintaining state-of-the-art performance on static and low-ego-motion datasets. Our key insight is to connect two powerful paradigms in modern 3D vision: the rich perceptual frontend of feedforward models and the end-to-end optimization of differentiable bundle adjustment (BA). We achieve this with a 3D-aware update operator built on a frozen, pre-trained MASt3R feature backbone, together with a high-capacity motion mask detector that uses multi-level 3D-aware features from the same backbone. Extensive experiments show WildPose consistently outperforms prior methods across dynamic (Wild-SLAM, Bonn), static (TUM, 7-Scenes), and low-ego-motion (Sintel) benchmarks.