Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: https://github.com/WangYuLin-SEU/WAPR.
Figures & tables
Figure 1 : Visualization of multi-iteration refinement for candidate poses with large angular deviations. The example is from TUD-L [ 18 ] , where the handheld object moves continuously. With large deviations between the candidate pose and the ground-truth pose, conventional small-angle refiners often fail, whereas WAPR corrects the pose through multi-iteration refinement.
Figure 2 : Overview of the proposed framework. Top: SA6D constructs large-scale RGB-D training data with KASAL-assisted rotational symmetry priors. These priors are used to canonicalize symmetry-equivalent training pairs. Bottom: WAPR refines sparse candidate poses by comparing real and rendered RGB-D observations with masks. The refined candidates are selected by a score-based pose selection module, enabling accurate 6D pose estimation under large initial misalignments of up to 90∘ .
Figure 3 : Large-scale construction of SA6D with rotational symmetry priors. KASAL-assisted annotation provides texture-aware and geometry-only priors for the original GSO scans. Geometry scaling and texture randomization reuse these priors to generate ∼ 20K non-symmetric and ∼ 30K rotationally symmetric augmented instances, and BlenderProc [ 6 ] renders ∼ 2M synthetic RGB-D images for training.
Figure 4 : Rotational symmetry ambiguity in wide-angle training. The candidate and target views can have a large raw rotation gap, while their visible difference may be limited to subtle local texture details, as shown by the zoomed regions. With rotational symmetry priors, such pairs are treated as symmetry-equivalent and canonicalized to the same supervision before loss computation.
Num
4
8
12
24
40
60
MSSD
59.6
67.3
71.4
71.9
72.8
73.3
MSPD
54.2
63.1
68.9
69.8
70.4
70.9
Mean
56.9
65.2
70.1
70.9
71.6
72.1
Table 1 : Analysis of candidate-pose coverage on IC-BIN. IC-BIN contains multiple identical objects in cluttered bin-like scenes, making it sensitive to the coverage of initial pose hypotheses. It is therefore used as a representative stress test for evaluating the number of candidate poses. “Num” denotes the number of candidate poses. Results are reported in AP for MSSD and MSPD, together with their mean.
Row
Method
IC-BIN [ 7 ]
TUD-L [ 18 ]
LM-O [ 1 ]
Mean
MSSD
MSPD
MSSD
MSPD
MSSD
MSPD
A0
WAPR(±90°)
71.9
69.8
96.0
95.4
77.0
80.7
81.8
B0
WAPR(±20°)
61.5
58.0
77.4
76.5
68.0
71.0
68.7
B1
WAPR(±40°)
69.9
67.1
95.0
94.7
74.4
77.8
79.8
B2
WAPR(±60°)
71.7
68.9
95.6
95.1
74.9
78.3
80.8
B3
WAPR(±140°)
70.6
68.1
95.5
95.0
74.4
77.7
80.2
Table 2 : Ablations of WAPR components on BOP detection AP. Results are reported as AP for MSSD and MSPD on three representative datasets, together with their mean. IC-BIN, TUD-L, and LM-O are selected to cover crowded multi-instance scenes, large object motion and viewpoint variation, and heavy occlusion, respectively. “w/o” denotes “without.” “WAPR ( ±90∘ )” is trained with candidate poses obtained by randomly perturbing the ground-truth pose within ±90∘ .
Method
LM-O
YCB-V
Mean
Time (s/img)
OSOP [ 48 ] + ICP
48.2
57.2
52.7
4
(PPF [ 10 ] , SIFT) + Zephyr [ 41 ]
59.8
51.6
55.7
–
MegaPose-RGBD [ 27 ]
58.3
63.3
60.8
–
FoundationPose [ 55 ] (MUSE) + 12 hypotheses
67.4
80.8
74.1
–
ICP (MUSE) + 12 hypotheses
26.6
29.7
28.2
35
WAPR (MUSE) + 12 hypotheses
75.5
89.3
82.4
1
Table 3 : Comparison with representative RGB-D refinement baselines on LM-O and YCB-V, where comparable prior-refinement results are available. The last three rows use MUSE detections and 12 hypotheses per instance.
Model-based 6D localization of unseen objects – BOP-Classic-Core
Method
2D Det.
LM-O
T-LESS
TUD-L
IC-BIN
ITODD
HB
YCB-V
Mean
Time (s)
fast
Co-op [ 37 ]
F3DT2D
73.0
68.0
92.9
62.4
60.0
86.3
88.6
75.9
0.765
WAPR*
F3DT2D
75.4
72.0
92.5
66.3
65.0
86.7
89.1
78.1
0.904
WAPR*
MUSE
75.4
73.1
96.1
70.9
70.5
90.1
88.3
80.6
0.72
unconstrained
MegaPose [ 27 ]
CNOS
62.6
48.7
85.1
46.7
46.8
73.0
76.4
62.8
141.965
Genflow [ 36 ]
CNOS
67.8
55.6
81.1
56.3
57.5
79.1
82.5
68.6
11.14
Table 4 : Comparison with SOTA methods on seven BOP-Classic-Core datasets for 6D localization (AR, top) and detection (AP, bottom). “2D Det.” reports the detector or hypothesis source. WAPR is trained on SA6D, while other methods use original training or released settings. “Multi-Det.” combines NIDS [ 35 ] , MUSE [ 5 ] , SAM6D [ 32 ] , and CNOS [ 38 ] . FP denotes FoundationPose [ 55 ] . “Time” reports per-image runtime, and “*” marks proposed methods.