cs.CVMar 27, 2024

Accurate and Efficient Object Pose Estimation via the Aggregation of Diffusion Features

Authors: Tianfu Wang, Guosheng Hu, Hongguang Wang

Organizations: State Key Laboratory of Robotics, Shenyang Institute of Automation, Chinese Academy of Sciences · Institutes for Robotics and Intelligent Manufacturing, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, 100049, China · Oosto, Belfast, U.K.

Abstract

Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To verify the efficacy of diffusion features for object pose estimation, we propose three distinct architectures (vanilla, nonlinear, and context-aware weight aggregations) that capture and aggregate diffusion features for comparative analysis. To achieve an efficient feature aggregation, we propose a confidence adaptive aggregation network that automatically selects the discriminative features rather than uses all the features, achieving a better speed-and-accuracy trade-off. In particular, our confidence adaptive aggregation network achieves higher accuracy than the previous best arts on unseen objects: 97.7% vs. 93.5% on Unseen LM, 85.5% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. On the large-scale BOP benchmark, our method also provides measurable gains, with an average recall of 58.3 compared to 57.9 previously. In addition, CAA reduces computational cost by 1.3-1.5 compared to the CWA variant while maintaining comparable accuracy. Furthermore, CAA reaches real-time performance, achieving over 68 FPS and offering a substantially improved accuracy-efficiency trade-off.

Figures & tables

Explore similar work

CardsList
  1. UniPose9D: Universal Category-Agnostic Object Pose Estimation

    Jul 10, 2026Yang You, Yi Du, Cole Harrison +1Object Pose Estimation

  2. WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation

    Oct 7, 2026Yulin Wang, Mengting Hu, Hongli Li +2Object Pose Estimation

  3. Learning Cross-View Semantic Priors for Single-Reference Unseen Object Pose Estimation

    Jun 20, 2026Jiahong Chen, Jinghao Wang, Ziwen Wang +3Multi-View ConsistencyObject Pose Estimation