Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: https://github.com/WangYuLin-SEU/WAPR.
Figures & tables
Figure 1 : Visualization of multi-iteration refinement for candidate poses with large angular deviations. The example is from TUD-L [ 18 ] , where the handheld object moves continuously. With large deviations between the candidate pose and the ground-truth pose, conventional small-angle refiners often fail, whereas WAPR corrects the pose through multi-iteration refinement.
Figure 2 : Overview of the proposed framework. Top: SA6D constructs large-scale RGB-D training data with KASAL-assisted rotational symmetry priors. These priors are used to canonicalize symmetry-equivalent training pairs. Bottom: WAPR refines sparse candidate poses by comparing real and rendered RGB-D observations with masks. The refined candidates are selected by a score-based pose selection module, enabling accurate 6D pose estimation under large initial misalignments of up to 90∘ .
Figure 3 : Large-scale construction of SA6D with rotational symmetry priors. KASAL-assisted annotation provides texture-aware and geometry-only priors for the original GSO scans. Geometry scaling and texture randomization reuse these priors to generate ∼ 20K non-symmetric and ∼ 30K rotationally symmetric augmented instances, and BlenderProc [ 6 ] renders ∼ 2M synthetic RGB-D images for training.
Figure 4 : Rotational symmetry ambiguity in wide-angle training. The candidate and target views can have a large raw rotation gap, while their visible difference may be limited to subtle local texture details, as shown by the zoomed regions. With rotational symmetry priors, such pairs are treated as symmetry-equivalent and canonicalized to the same supervision before loss computation.
Num
4
8
12
24
40
60
MSSD
59.6
67.3
71.4
71.9
72.8
73.3
MSPD
54.2
63.1
68.9
69.8
70.4
70.9
Mean
56.9
65.2
70.1
70.9
71.6
72.1
Table 1 : Analysis of candidate-pose coverage on IC-BIN. IC-BIN contains multiple identical objects in cluttered bin-like scenes, making it sensitive to the coverage of initial pose hypotheses. It is therefore used as a representative stress test for evaluating the number of candidate poses. “Num” denotes the number of candidate poses. Results are reported in AP for MSSD and MSPD, together with their mean.
Row
Method
IC-BIN [ 7 ]
TUD-L [ 18 ]
LM-O [ 1 ]
Mean
MSSD
MSPD
MSSD
MSPD
MSSD
MSPD
A0
WAPR(±90°)
71.9
69.8
96.0
95.4
77.0
80.7
81.8
B0
WAPR(±20°)
61.5
58.0
77.4
76.5
68.0
71.0
68.7
B1
WAPR(±40°)
69.9
67.1
95.0
94.7
74.4
77.8
79.8
B2
WAPR(±60°)
71.7
68.9
95.6
95.1
74.9
78.3
80.8
B3
WAPR(±140°)
70.6
68.1
95.5
95.0
74.4
77.7
80.2
Table 2 : Ablations of WAPR components on BOP detection AP. Results are reported as AP for MSSD and MSPD on three representative datasets, together with their mean. IC-BIN, TUD-L, and LM-O are selected to cover crowded multi-instance scenes, large object motion and viewpoint variation, and heavy occlusion, respectively. “w/o” denotes “without.” “WAPR ( ±90∘ )” is trained with candidate poses obtained by randomly perturbing the ground-truth pose within ±90∘ .
Method
LM-O
YCB-V
Mean
Time (s/img)
OSOP [ 48 ] + ICP
48.2
57.2
52.7
4
(PPF [ 10 ] , SIFT) + Zephyr [ 41 ]
59.8
51.6
55.7
–
MegaPose-RGBD [ 27 ]
58.3
63.3
60.8
–
FoundationPose [ 55 ] (MUSE) + 12 hypotheses
67.4
80.8
74.1
–
ICP (MUSE) + 12 hypotheses
26.6
29.7
28.2
35
WAPR (MUSE) + 12 hypotheses
75.5
89.3
82.4
1
Table 3 : Comparison with representative RGB-D refinement baselines on LM-O and YCB-V, where comparable prior-refinement results are available. The last three rows use MUSE detections and 12 hypotheses per instance.
Model-based 6D localization of unseen objects – BOP-Classic-Core
Method
2D Det.
LM-O
T-LESS
TUD-L
IC-BIN
ITODD
HB
YCB-V
Mean
Time (s)
fast
Co-op [ 37 ]
F3DT2D
73.0
68.0
92.9
62.4
60.0
86.3
88.6
75.9
0.765
WAPR*
F3DT2D
75.4
72.0
92.5
66.3
65.0
86.7
89.1
78.1
0.904
WAPR*
MUSE
75.4
73.1
96.1
70.9
70.5
90.1
88.3
80.6
0.72
unconstrained
MegaPose [ 27 ]
CNOS
62.6
48.7
85.1
46.7
46.8
73.0
76.4
62.8
141.965
Genflow [ 36 ]
CNOS
67.8
55.6
81.1
56.3
57.5
79.1
82.5
68.6
11.14
Table 4 : Comparison with SOTA methods on seven BOP-Classic-Core datasets for 6D localization (AR, top) and detection (AP, bottom). “2D Det.” reports the detector or hypothesis source. WAPR is trained on SA6D, while other methods use original training or released settings. “Multi-Det.” combines NIDS [ 35 ] , MUSE [ 5 ] , SAM6D [ 32 ] , and CNOS [ 38 ] . FP denotes FoundationPose [ 55 ] . “Time” reports per-image runtime, and “*” marks proposed methods.
Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving from explicit 3D models to multi-view object captures to single reference images. We take this progression to its extreme by introducing prior-free relative 6D pose estimation, which lifts the assumption of knowing which object is to be posed within the scene. This novel setting aims to estimate the relative poses of multiple instances of an unknown object within the same image, without requiring CAD models, templates, or reference images. We solve this by formulating a novel method (PROSE) that finds coarse correspondences between object instances using multimodal foundation features, thus requiring no training. We refine these correspondences by imposing cycle consistency across tuples of instances, and leverage the resulting globally consistent correspondences to estimate the relative 6D pose between any pair of instances. To enable systematic evaluation, we design a novel benchmark (PRENCH) built from three multi-instance BOP datasets and enriched with task-specific metadata. PROSE consistently outperforms baselines obtained by adapting state-of-the-art single-image methods to the proposed setting, while requiring neither task-specific supervision nor additional learned components. Project website: https://tev-fbk.github.io/PROSE/
Behdad Khodabandehloo, Andrea Caraffa, Davide Boscaini +1
Fondazione Bruno Kessler, Trento, Italy · University of Trento, Trento, Italy
In many practical 6D object pose estimation scenarios, we often have access to only a single real-world RGB-D reference view per object, typically without CAD models. Existing methods largely rely on explicit 3D models or multi-view data, which limits their scalability. To address this challenging single-reference model-free setting, we propose \textbf{OneViewAll}, a semantic-prior-guided framework that performs pose estimation via a novel Project-and-Compare paradigm. Instead of relying on computationally expensive CAD-based rendering, our method directly aligns reference and query observations within a projection-equivariant space. OneViewAll progressively integrates hierarchical semantic priors across three levels: (1) \textit{category- and scene-level} priors for efficient hypothesis initialization; (2) \textit{object-level symmetry} priors for geometry completion via mirror fusion; and (3) \textit{patch-level} priors for discriminative refinement. Extensive experiments demonstrate that OneViewAll achieves \textbf{92.5%} ADD-0.1 accuracy on the LINEMOD dataset using only one real reference view -- significantly outperforming the CVPR 2025 baseline One2Any (52.6%). It also yields consistent improvements on YCB-V, Real275, and Toyota-Light while maintaining low inference latency. Our results underscore the efficacy of symmetry-aware projection in handling symmetric, texture-less, and occluded objects.
Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, generalization to novel categories and unseen scenes remains challenging. We propose UniPose9D, a unified model for category-agnostic 9D object pose estimation: given an instance mask/ROI and either an RGB-D observation or an RGB image with predicted depth, the model estimates rotation, translation, and metric size without category labels, CAD models, mean-shape priors, or reference views. Specifically, UniPose9D samples point pairs from the observed object geometry and uses DINOv2 and PointNet features to predict NOCS coordinates for each pair. To improve accuracy, we introduce a point-pair-based RANSAC N-hop Kabsch-Umeyama algorithm with an adaptive threshold. We further employ flow matching to address symmetric ambiguities and construct a large-scale training set by curating and aligning pose annotations from existing public datasets. Experiments across eight datasets show that a single unified model achieves competitive performance on standard benchmarks while generalizing to unseen objects, unseen categories, and in-the-wild scenarios. Our code and model are available at https://github.com/qq456cvb/UniPose9D.