Accurate and Efficient Object Pose Estimation via the Aggregation of Diffusion Features
Authors: Tianfu Wang, Guosheng Hu, Hongguang Wang
Organizations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences · Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, 100049, China · Oosto, Belfast, U.K.
Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To verify the efficacy of diffusion features for object pose estimation, we propose three distinct architectures (vanilla, nonlinear, and context-aware weight aggregations) that capture and aggregate diffusion features for comparative analysis. To achieve an efficient feature aggregation, we propose a confidence adaptive aggregation network that automatically selects the discriminative features rather than uses all the features, achieving a better speed-and-accuracy trade-off. In particular, our confidence adaptive aggregation network achieves higher accuracy than the previous best arts on unseen objects: 97.7% vs. 93.5% on Unseen LM, 85.5% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. On the large-scale BOP benchmark, our method also provides measurable gains, with an average recall of 58.3 compared to 57.9 previously. In addition, CAA reduces computational cost by 1.3-1.5 compared to the CWA variant while maintaining comparable accuracy. Furthermore, CAA reaches real-time performance, achieving over 68 FPS and offering a substantially improved accuracy-efficiency trade-off.
Figures & tables
Figure 1: Unseen object pose estimation of one state-of-the-art method ( Nguyen et al., 2022 ) and our method. (a) We render an image using the ground truth pose of an unseen object. (b) Template-pose ( Nguyen et al., 2022 ) learns image features by fine-tuning a self-supervised learning ( He et al., 2020 ) pre-trained model. (c) Our method aggregrates features of different granularity from a diffusion model to achieve better pose estimation than template-pose ( Nguyen et al., 2022 ) on the unseen object. The snowflake and flame symbols represent ‘parameters frozen’ and ‘fine-tune’, respectively.
Figure 2: Feature visualization of LINEMOD. For query and template images, we visualize their 3 features from template-pose ( Nguyen et al., 2022 ) , Layer 5 and Layer 12 of a diffusion model. These features are projected to a PCA space, and the values of top 3 principal components are assigned to RGB values respectively for visualization. The more similar the colors of two feature images in one column, the more similar in feature space. The two features in one green box are very similar.
Figure 3: Diffusion aggregation methods aggregate all intermediate features.
Figure 4: Confidence Adaptive Aggregation (CAA) method aggregates features within and across blocks, with early exit decisions guided by confidence prediction modules.
Method
Backbone
Res.
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
MACs
FPS
Balntas et al. ( Balntas et al., 2017 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
87.0
13.2
83.1
15.5
85.1
18.2
85.0
15.2
11.9M
493
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
89.2
14.1
85.4
16.3
83.3
16.7
86.3
16.7
11.9M
493
MPL ( Sundermeyer et al., 2020 )
MPL ( Sundermeyer et al., 2020 )
1282
91.5
36.0
87.9
39.3
87.2
39.7
88.9
38.3
2.1B
52.4
Zhao et al. ( Zhao et al., 2022 )
ResNet ( He et al., 2016 )
1282
96.9
87.5
94.4
76.2
93.1
73.7
93.8
79.2
1.7B
43.8
Template-pose ( Nguyen et al., 2022 )
ResNet ( He et al., 2016 )
5122
99.3
95.8
99.1
97.9
99.4
89.6
99.3
94.4
130B
46.5
Table 1: Results on LM ( Hinterstoisser et al., 2013 ) . VA represents vanilla aggregation, NA represents nonlinear aggregation, CWA represents context-aware weight aggregation, and CAA represents confidence adaptive aggregation.
Method
Backbone
Res.
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
MACs
FPS
Balntas et al. ( Balntas et al., 2017 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
19.2
9.3
23.1
5.1
15.0
5.1
19.1
6.5
11.9M
493
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
18.3
8.2
21.9
7.5
17.6
7.6
19.5
7.8
11.9M
493
MPL ( Sundermeyer et al., 2020 )
MPL ( Sundermeyer et al., 2020 )
1282
31.3
18.6
34.5
15.9
29.2
17.7
31.7
17.4
2.1B
52.4
Zhao et al. ( Zhao et al., 2022 )
ResNet ( He et al., 2016 )
1282
54.9
40.1
63.4
32.7
49.9
37.5
56.1
36.8
1.7B
43.8
Template-pose ( Nguyen et al., 2022 )
ResNet ( He et al., 2016 )
5122
78.5
73.6
85.2
73.1
78.0
87.2
80.6
77.9
130B
46.5
Table 2: Results on O-LM ( Brachmann et al., 2014 ) . VA represents vanilla aggregation, NA represents nonlinear aggregation, CWA represents context-aware weight aggregation, and CAA represents confidence adaptive aggregation.
Figure 5: Comparison of accuracy-efficiency trade-offs across methods on the unseen O-LM, illustrating the relationship between inference speed (FPS) and pose estimation accuracy. Numbers in parentheses indicate the input resolution used for each method.
Method
Number templates
Res.
Recall VSD
MACs
Obj. 1-18
Obj. 19-30
Avg.
Implicit ( Sundermeyer et al., 2018 )
92K
1282
35.60
42.45
38.34
2.1B
MPL ( Sundermeyer et al., 2020 )
92K
1282
35.25
33.17
34.42
2.1B
Template-pose ( Nguyen et al., 2022 )
92K
2242
59.62
57.75
58.87
24.9B
Template-pose ( Nguyen et al., 2022 )
21K
5122
61.11
57.97
59.59
130B
Template-pose ( Nguyen et al., 2022 )
21K
2242
59.14
56.91
58.25
24.9B
Table 3: Results on T-LESS ( Hodan et al., 2017 ) .
Method
Pose refine.
Trainable Params
LM-O
T-LESS
TUD-L
ICBIN
ITODD
HB
YCB-V
Avg.
Time
Pose estimation (single hypothesis)
MegaPose ( Labbé et al., 2022 )
MegaPose ( Labbé et al., 2022 )
64.9 M
49.9
47.7
65.3
36.7
31.5
65.4
60.1
50.9
31.7 s
FoundPose ( Örnek et al., 2024 )
MegaPose ( Labbé et al., 2022 )
43.1 M
55.4
51.0
63.3
43.0
34.6
69.5
66.1
54.7
4.4 s
GigaPose ( Nguyen et al., 2024 )
MegaPose ( Labbé et al., 2022 )
360 M
55.6
54.6
57.8
44.3
37.8
69.6
63.4
54.7
2.4 s
Ours
MegaPose ( Labbé et al., 2022 )
70.7 M
56.5
53.6
62.2
45.1
36.0
70.3
64.6
55.5
2.2 s
Pose estimation (five hypotheses)
Table 4: Comparison with state-of-the-art methods on the seven core BOP datasets. We report Average Recall (AR) under the evaluation settings defined in the BOP Challenge, including both single- and multi-hypothesis predictions.
Figure 6: Ablation on timestep t . The accuracy is measured for the model by extracting features from Stable Diffusion at different timesteps on LM and O-LM datasets.
Backbone
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
SD-V1-5 ( Rombach et al., 2022 )
99.5
98.3
99.6
98.9
99.8
96.3
99.7
97.9
SD-V2-0 ( Rombach et al., 2022 )
99.6
98.1
99.4
99.5
99.6
97.5
99.5
98.4
OpenCLIP ( Radford et al., 2021b )
99.4
97.0
99.7
98.5
99.7
95.9
99.6
97.1
DINOv2 ( Oquab et al., 2023 )
99.7
96.6
99.6
99.1
99.5
97.8
99.6
97.8
Table 5: Comparison with other large models on LM ( Hinterstoisser et al., 2013 ) .
Backbone
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
SD-V1-5 ( Rombach et al., 2022 )
81.5
81.6
86.0
79.8
80.8
96.3
82.8
85.9
SD-V2-0 ( Rombach et al., 2022 )
82.9
81.0
88.1
79.1
79.2
95.5
83.4
85.2
OpenCLIP ( Radford et al., 2021b )
79.8
76.5
85.7
73.9
80.1
91.9
81.9
81.1
DINOv2 ( Oquab et al., 2023 )
84.1
73.7
96.1
74.3
79.0
86.3
83.1
78.1
Table 6: Comparison with other large models on O-LM ( Brachmann et al., 2014 ) .
δ
Avg. on LM
Avg. on O-LM
Avg.
Seen
Unseen
Seen
Unseen
MACs
FPS
0.65
99.4
97.1
82.5
83.0
610B
17.6
0.7
99.6
97.7
83.5
85.5
719B
14.9
0.75
99.7
97.9
83.6
85.6
863B
12.4
Table 7: Impact of confidence threshold δ .
Figure 7: Visualization of the impact of varying confidence thresholds δ .
Exit Block
LM Seen
LM Unseen
2
0.743
0.694
4
0.752
0.701
6
0.736
0.683
8
0.759
0.711
Table 8: AUROC of the CAA confidence score for discriminating correct and incorrect retrievals at different exit blocks on LM. Higher is better.
Subset
ECE ↓
Brier Score ↓
LM Seen
0.006
0.0101
LM Unseen
0.026
0.0229
Table 9: Empirical calibration of the CAA confidence score on LM using the final-block confidence of Splits #2 and #3. Lower is better.
Figure 8: Exit-depth distributions of CAA under δ=0.7 on LM and O-LM seen objects.
No. of agg. features
2
4
6
Fixed budget
82.0
85.0
85.3
CAA*
83.2
85.5
85.7
CAA* − Fixed (pp)
+1.2±0.3
+0.5±0.1
+0.4±0.1
Table 10: Impact of confidence adaptive exiting. Results are averaged over three object splits. For each column, CAA* operates with the corresponding average number of aggregated features. The ± values denote the sample standard deviation of the paired CAA–Fixed (pp) accuracy differences across the three splits.
Feature
Aggregation
Single-hyp. AR
Five-hyp. AR
DINOv2
Standard
54.7
57.9
DINOv2
CAA
54.5
57.8
Diffusion
Standard
50.2
53.1
Diffusion
CAA
55.5
58.3
Table 11: Controlled ablation of feature representation and aggregation strategy under the BOP setting.
Figure 9: UMAP visualizations of feature representations after removing backgrounds using ground-truth object masks.
Figure 10: Correlation between angle difference and feature similarity for template-pose (top) and our method (bottom) on the unseen objects of LM split #1 (ape, bench, can, cam). The shaded region (0-25°) highlights small pose differences, with the Kendall-tau coefficient used to assess pose discriminability.
Figure 11: Qualitative results on unseen objects of O-LM (left) and T-LESS (right).
Template-pose
VA
NA
CWA
CAA
Trainable Params (M)
24.04
2.53
13.28
13.29
13.27
Table 12: Comparison of trainable parameter sizes for different methods.
Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, generalization to novel categories and unseen scenes remains challenging. We propose UniPose9D, a unified model for category-agnostic 9D object pose estimation: given an instance mask/ROI and either an RGB-D observation or an RGB image with predicted depth, the model estimates rotation, translation, and metric size without category labels, CAD models, mean-shape priors, or reference views. Specifically, UniPose9D samples point pairs from the observed object geometry and uses DINOv2 and PointNet features to predict NOCS coordinates for each pair. To improve accuracy, we introduce a point-pair-based RANSAC N-hop Kabsch-Umeyama algorithm with an adaptive threshold. We further employ flow matching to address symmetric ambiguities and construct a large-scale training set by curating and aligning pose annotations from existing public datasets. Experiments across eight datasets show that a single unified model achieves competitive performance on standard benchmarks while generalizing to unseen objects, unseen categories, and in-the-wild scenarios. Our code and model are available at https://github.com/qq456cvb/UniPose9D.
Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: https://github.com/WangYuLin-SEU/WAPR.
Yulin Wang, Mengting Hu, Hongli Li +2
Southeast University, Nanjing, China · Purdue University, West Lafayette, IN, USA
Single-reference unseen object 6D pose estimation reduces object onboarding by estimating poses of arbitrary novel objects from only one reference view. Recent correspondence-based pipelines have achieved robust performance with vision foundation model (VFM) features. However, they typically treat these features as intra-view descriptors, leaving dense visual-semantic cues, including appearance, structure, and context, insufficiently exchanged across views before geometric decoding. Consequently, the decoded point features may lack joint semantic and geometric discriminability, making correspondence estimation still difficult in challenging cases. Instead of processing features independently, we build the correspondence pipeline around an early cross-view semantic prior. Specifically, cross-view semantic interaction (CVSI) enables dense query and reference VFM tokens to exchange semantic context and form a cross-view prior. Nevertheless, direct CVSI may disturb the VFM token structure, while the resulting semantic prior still needs 3D representation consistency for rigid correspondence. To make this CVSI prior reliable for 3D correspondence learning, we introduce two complementary training-time constraints: the intra-view structure preservation (IVSP) loss preserves the original intra-view token affinity structure during interaction, while the reference-anchored geometric consistency (RAGC) loss enforces spatial representation consistency of decoded point features. The final pose is recovered from learned correspondences through weighted SVD. We further construct a challenging view-pair protocol from the BOP Challenge datasets YCB-V and TUD-L to evaluate robustness in difficult matching scenarios. Extensive experiments on six benchmarks under different view-pair settings show that our method achieves state-of-the-art performance while maintaining comparable inference speed.
Jiahong Chen, Jinghao Wang, Ziwen Wang +3
College of Aerospace Science and Engineering, National University of Defense Technology, Changsha 410073, China · Hunan Provincial Key Laboratory of Image Measurement and Vision Navigation, Changsha 410073, China