Accurate and Efficient Object Pose Estimation via the Aggregation of Diffusion Features
Authors: Tianfu Wang, Guosheng Hu, Hongguang Wang
Organizations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences · Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, 100049, China · Oosto, Belfast, U.K.
Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To verify the efficacy of diffusion features for object pose estimation, we propose three distinct architectures (vanilla, nonlinear, and context-aware weight aggregations) that capture and aggregate diffusion features for comparative analysis. To achieve an efficient feature aggregation, we propose a confidence adaptive aggregation network that automatically selects the discriminative features rather than uses all the features, achieving a better speed-and-accuracy trade-off. In particular, our confidence adaptive aggregation network achieves higher accuracy than the previous best arts on unseen objects: 97.7% vs. 93.5% on Unseen LM, 85.5% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. On the large-scale BOP benchmark, our method also provides measurable gains, with an average recall of 58.3 compared to 57.9 previously. In addition, CAA reduces computational cost by 1.3-1.5 compared to the CWA variant while maintaining comparable accuracy. Furthermore, CAA reaches real-time performance, achieving over 68 FPS and offering a substantially improved accuracy-efficiency trade-off.
Figures & tables
Figure 1: Unseen object pose estimation of one state-of-the-art method ( Nguyen et al., 2022 ) and our method. (a) We render an image using the ground truth pose of an unseen object. (b) Template-pose ( Nguyen et al., 2022 ) learns image features by fine-tuning a self-supervised learning ( He et al., 2020 ) pre-trained model. (c) Our method aggregrates features of different granularity from a diffusion model to achieve better pose estimation than template-pose ( Nguyen et al., 2022 ) on the unseen object. The snowflake and flame symbols represent ‘parameters frozen’ and ‘fine-tune’, respectively.
Figure 2: Feature visualization of LINEMOD. For query and template images, we visualize their 3 features from template-pose ( Nguyen et al., 2022 ) , Layer 5 and Layer 12 of a diffusion model. These features are projected to a PCA space, and the values of top 3 principal components are assigned to RGB values respectively for visualization. The more similar the colors of two feature images in one column, the more similar in feature space. The two features in one green box are very similar.
Figure 3: Diffusion aggregation methods aggregate all intermediate features.
Figure 4: Confidence Adaptive Aggregation (CAA) method aggregates features within and across blocks, with early exit decisions guided by confidence prediction modules.
Method
Backbone
Res.
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
MACs
FPS
Balntas et al. ( Balntas et al., 2017 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
87.0
13.2
83.1
15.5
85.1
18.2
85.0
15.2
11.9M
493
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
89.2
14.1
85.4
16.3
83.3
16.7
86.3
16.7
11.9M
493
MPL ( Sundermeyer et al., 2020 )
MPL ( Sundermeyer et al., 2020 )
1282
91.5
36.0
87.9
39.3
87.2
39.7
88.9
38.3
2.1B
52.4
Zhao et al. ( Zhao et al., 2022 )
ResNet ( He et al., 2016 )
1282
96.9
87.5
94.4
76.2
93.1
73.7
93.8
79.2
1.7B
43.8
Template-pose ( Nguyen et al., 2022 )
ResNet ( He et al., 2016 )
5122
99.3
95.8
99.1
97.9
99.4
89.6
99.3
94.4
130B
46.5
Table 1: Results on LM ( Hinterstoisser et al., 2013 ) . VA represents vanilla aggregation, NA represents nonlinear aggregation, CWA represents context-aware weight aggregation, and CAA represents confidence adaptive aggregation.
Method
Backbone
Res.
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
MACs
FPS
Balntas et al. ( Balntas et al., 2017 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
19.2
9.3
23.1
5.1
15.0
5.1
19.1
6.5
11.9M
493
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
Wohlhart et al. ( Wohlhart and Lepetit, 2015 )
642
18.3
8.2
21.9
7.5
17.6
7.6
19.5
7.8
11.9M
493
MPL ( Sundermeyer et al., 2020 )
MPL ( Sundermeyer et al., 2020 )
1282
31.3
18.6
34.5
15.9
29.2
17.7
31.7
17.4
2.1B
52.4
Zhao et al. ( Zhao et al., 2022 )
ResNet ( He et al., 2016 )
1282
54.9
40.1
63.4
32.7
49.9
37.5
56.1
36.8
1.7B
43.8
Template-pose ( Nguyen et al., 2022 )
ResNet ( He et al., 2016 )
5122
78.5
73.6
85.2
73.1
78.0
87.2
80.6
77.9
130B
46.5
Table 2: Results on O-LM ( Brachmann et al., 2014 ) . VA represents vanilla aggregation, NA represents nonlinear aggregation, CWA represents context-aware weight aggregation, and CAA represents confidence adaptive aggregation.
Figure 5: Comparison of accuracy-efficiency trade-offs across methods on the unseen O-LM, illustrating the relationship between inference speed (FPS) and pose estimation accuracy. Numbers in parentheses indicate the input resolution used for each method.
Method
Number templates
Res.
Recall VSD
MACs
Obj. 1-18
Obj. 19-30
Avg.
Implicit ( Sundermeyer et al., 2018 )
92K
1282
35.60
42.45
38.34
2.1B
MPL ( Sundermeyer et al., 2020 )
92K
1282
35.25
33.17
34.42
2.1B
Template-pose ( Nguyen et al., 2022 )
92K
2242
59.62
57.75
58.87
24.9B
Template-pose ( Nguyen et al., 2022 )
21K
5122
61.11
57.97
59.59
130B
Template-pose ( Nguyen et al., 2022 )
21K
2242
59.14
56.91
58.25
24.9B
Table 3: Results on T-LESS ( Hodan et al., 2017 ) .
Method
Pose refine.
Trainable Params
LM-O
T-LESS
TUD-L
ICBIN
ITODD
HB
YCB-V
Avg.
Time
Pose estimation (single hypothesis)
MegaPose ( Labbé et al., 2022 )
MegaPose ( Labbé et al., 2022 )
64.9 M
49.9
47.7
65.3
36.7
31.5
65.4
60.1
50.9
31.7 s
FoundPose ( Örnek et al., 2024 )
MegaPose ( Labbé et al., 2022 )
43.1 M
55.4
51.0
63.3
43.0
34.6
69.5
66.1
54.7
4.4 s
GigaPose ( Nguyen et al., 2024 )
MegaPose ( Labbé et al., 2022 )
360 M
55.6
54.6
57.8
44.3
37.8
69.6
63.4
54.7
2.4 s
Ours
MegaPose ( Labbé et al., 2022 )
70.7 M
56.5
53.6
62.2
45.1
36.0
70.3
64.6
55.5
2.2 s
Pose estimation (five hypotheses)
Table 4: Comparison with state-of-the-art methods on the seven core BOP datasets. We report Average Recall (AR) under the evaluation settings defined in the BOP Challenge, including both single- and multi-hypothesis predictions.
Figure 6: Ablation on timestep t . The accuracy is measured for the model by extracting features from Stable Diffusion at different timesteps on LM and O-LM datasets.
Backbone
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
SD-V1-5 ( Rombach et al., 2022 )
99.5
98.3
99.6
98.9
99.8
96.3
99.7
97.9
SD-V2-0 ( Rombach et al., 2022 )
99.6
98.1
99.4
99.5
99.6
97.5
99.5
98.4
OpenCLIP ( Radford et al., 2021b )
99.4
97.0
99.7
98.5
99.7
95.9
99.6
97.1
DINOv2 ( Oquab et al., 2023 )
99.7
96.6
99.6
99.1
99.5
97.8
99.6
97.8
Table 5: Comparison with other large models on LM ( Hinterstoisser et al., 2013 ) .
Backbone
Split #1
Split #2
Split #3
Avg.
Seen
Unseen
Seen
Unseen
Seen
Unseen
Seen
Unseen
SD-V1-5 ( Rombach et al., 2022 )
81.5
81.6
86.0
79.8
80.8
96.3
82.8
85.9
SD-V2-0 ( Rombach et al., 2022 )
82.9
81.0
88.1
79.1
79.2
95.5
83.4
85.2
OpenCLIP ( Radford et al., 2021b )
79.8
76.5
85.7
73.9
80.1
91.9
81.9
81.1
DINOv2 ( Oquab et al., 2023 )
84.1
73.7
96.1
74.3
79.0
86.3
83.1
78.1
Table 6: Comparison with other large models on O-LM ( Brachmann et al., 2014 ) .
δ
Avg. on LM
Avg. on O-LM
Avg.
Seen
Unseen
Seen
Unseen
MACs
FPS
0.65
99.4
97.1
82.5
83.0
610B
17.6
0.7
99.6
97.7
83.5
85.5
719B
14.9
0.75
99.7
97.9
83.6
85.6
863B
12.4
Table 7: Impact of confidence threshold δ .
Figure 7: Visualization of the impact of varying confidence thresholds δ .
Exit Block
LM Seen
LM Unseen
2
0.743
0.694
4
0.752
0.701
6
0.736
0.683
8
0.759
0.711
Table 8: AUROC of the CAA confidence score for discriminating correct and incorrect retrievals at different exit blocks on LM. Higher is better.
Subset
ECE ↓
Brier Score ↓
LM Seen
0.006
0.0101
LM Unseen
0.026
0.0229
Table 9: Empirical calibration of the CAA confidence score on LM using the final-block confidence of Splits #2 and #3. Lower is better.
Figure 8: Exit-depth distributions of CAA under δ=0.7 on LM and O-LM seen objects.
No. of agg. features
2
4
6
Fixed budget
82.0
85.0
85.3
CAA*
83.2
85.5
85.7
CAA* − Fixed (pp)
+1.2±0.3
+0.5±0.1
+0.4±0.1
Table 10: Impact of confidence adaptive exiting. Results are averaged over three object splits. For each column, CAA* operates with the corresponding average number of aggregated features. The ± values denote the sample standard deviation of the paired CAA–Fixed (pp) accuracy differences across the three splits.
Feature
Aggregation
Single-hyp. AR
Five-hyp. AR
DINOv2
Standard
54.7
57.9
DINOv2
CAA
54.5
57.8
Diffusion
Standard
50.2
53.1
Diffusion
CAA
55.5
58.3
Table 11: Controlled ablation of feature representation and aggregation strategy under the BOP setting.
Figure 9: UMAP visualizations of feature representations after removing backgrounds using ground-truth object masks.
Figure 10: Correlation between angle difference and feature similarity for template-pose (top) and our method (bottom) on the unseen objects of LM split #1 (ape, bench, can, cam). The shaded region (0-25°) highlights small pose differences, with the Kendall-tau coefficient used to assess pose discriminability.
Figure 11: Qualitative results on unseen objects of O-LM (left) and T-LESS (right).
Template-pose
VA
NA
CWA
CAA
Trainable Params (M)
24.04
2.53
13.28
13.29
13.27
Table 12: Comparison of trainable parameter sizes for different methods.
College of Aerospace Science and Engineering, National University of Defense Technology, Changsha 410073, China · Hunan Provincial Key Laboratory of Image Measurement and Vision Navigation, Changsha 410073, China