LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation
Authors: Hongli Xu, Zhaowei Lu, Junwen Huang, Jiaqi Hu, Peter KT Yu, Benjamin Busam, Federico Tombari, Slobodan ilic
Organizations: Technical University of Munich Munich, Germany · Munich Center for Machine Learning Munich, Germany · Siemens AG Munich, Germany · ROBOX Cambridge, MA, USA
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object pose and size, and a canonical Semantic Gaussian Field. Rather than treating reconstruction as a detached auxiliary task, LEGAU uses the Gaussian field as a category-conditioned structural prior that participates in multimodal feature fusion and provides global guidance for local pose reasoning. Conditioned on a categorical text embedding, LEGAU processes RGB-D observations through a transformer-based fusion module that integrates visual, geometric, and category-level cues, decoding the NOCS map, pose and size information and the Gaussian-based object representation. Extensive experiments on synthetic and real-world benchmarks show that this coupled pose-shape formulation achieves strong performance in a single-model multi-category setting, with up to 22% on SOPE and competitive transfer to real-world data. These results highlight the benefit of jointly learning canonical correspondence, object shape, and pose alignment within a unified representation.
Figures & tables
Figure 1: We present LEGAU , a unified framework for scalable category-level object pose estimation from a single view. Our method learns a Semantic Gaussian Field that provides shape- and category-aware guidance for coordinate alignment, enabling simultaneous object understanding and pose estimation, which may support downstream robotic perception tasks.
Figure 2: Pipeline of the LEGAU framework. Given an RGB-D input and a categorical text prompt, LEGAU extracts modality-specific features using DINOv2 [ 28 ] , CLIP [ 32 ] , and PointNet [ 31 ] , producing a text embedding that encodes category priors and a local embedding combining visual and geometric cues. A learnable field embedding captures global semantic and geometric priors and jointly attends with the multi-modal embeddings through a transformer. This unified representation enables coupled semantic shape understanding and pose reasoning: local embeddings drive the NOCS and pose decoders, while the global embedding guides the shape decoder to reconstruct high-level semantics and geometry as Gaussian primitives. The training of our model uses multi-level supervision, including NOCS and pose losses, photometric consistency, and a cosine-similarity objective for feature coherence.
Figure 3: Qualitative comparison between our method, AGPose [ 23 ] and GenPose++ [ 47 ] on SOPE, ROPE, and HouseCat6D. From left to right: a symmetric and occluded bowl (SOPE), a transparent glass (SOPE), a real-world toy object (ROPE), and reflective metallic tableware (HouseCat6D). These examples illustrate robustness to symmetry, occlusion, transparency, real-world novel objects, and specular surfaces.
Dataset
Method
IoU25
IoU50
IoU75
5°2cm
5°5cm
10°2cm
10°5cm
HouseCat6D [ 14 ]
VI-Net [ 6 ]
-
56.4
-
8.4
10.3
20.5
29.1
SecondPose [ 5 ]
-
66.1
-
11.0
13.4
25.3
35.7
AG-Pose [ 23 ]
88.1
76.9
53.0
21.3
22.1
51.3
54.3
LEGAU
90.7
80.6
57.4
21.4
22.7
54.0
57.4
Dataset
Method
AUC ↑ (IoU)
VUS ↑ ( n ° m cm)
IoU25
IoU50
IoU75
5°2cm
5°5cm
10°2cm
10°5cm
Table 1: Quantitative comparison on HouseCat6D [ 14 ] , SOPE [ 47 ] , and ROPE [ 47 ] . For each column, the top three methods are highlighted using a blue color map, where darker shades indicate better performance. LEGAU achieves strong performance on HouseCat6D and SOPE under the same evaluation protocol, especially in the single-model multi-category setting. On ROPE, LEGAU remains competitive in IoU-based metrics but shows a gap under stricter pose thresholds, indicating that synthetic-to-real transfer remains challenging..
Figure 4: Visual of Semantic Gaussians Reconstruction.It renders Gaussian primitives decoded in canonical object space, where the per-Gaussian semantic feature embeddings are projected to RGB colors using PCA.
GF
Sem. GF
10°5cm
IoU50
✓
✓
57.4
80.6
✓
✗
55.1 ( -4.4% )
78.3( -2.9% )
✗
✗
52.0( -9.4% )
75.0 ( -7.0% )
Table 2: Ablations on isolating the effects of Gaussian reconstruction (left) and CLIP-based category conditioning (right).
Figure 5: Qualitative comparison of shape reconstruction from partial point observations. We compare our method with a depth-based reconstruction baseline [ 45 ] and recent image-to-3D generative frameworks [ 40 , 39 , 42 ] under challenging scenarios involving heavy occlusion and transparent materials.
FoldingNet [ 41 ]
PoinTr [ 44 ]
AdaPoinTr [ 45 ]
Ours
62.72
29.87
23.17
6.72 ( -16.5% )
Table 3: Shape reconstruction results on the SOPE dataset. We evaluate on the Chamfer-L1 ( ×10−3 m, lower is better ). Best results are bolded .
Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, they often overfit to existing benchmarks and exhibit limited generalization to novel categories and unseen scenes. We propose UniPose9D, a category-agnostic foundation model for 9D object pose estimation: given an instance mask/ROI and either an RGB-D observation or an RGB image with predicted depth, the model estimates rotation, translation, and metric size without category labels, CAD models, mean-shape priors, or reference views. Specifically, UniPose9D samples point pairs from the observed object geometry and uses DINOv2 and PointNet features to predict NOCS coordinates for each pair. To improve accuracy, we introduce a point-pair-based RANSAC N-hop Kabsch--Umeyama algorithm with an adaptive threshold. We further employ flow matching to address symmetric ambiguities and construct a large-scale training set by curating and aligning pose annotations from existing public datasets. Experiments across six datasets show that a single unified model can match or surpass specialist methods while generalizing to unseen objects and in-the-wild scenarios. Our code and model are available on https://github.com/qq456cvb/UniPose9D.
Category-level object pose estimation seeks to recover a similarity transform (R,t,s) for unseen instances without instance-specific CAD models. Most competitive methods are correspondence-based: prior-free variants regress canonical (NOCS) coordinates directly from local observations and implicitly memorize the canonical frame in the weights, which ties the parameters to category-typical orientations and hurts generalization under distribution shift; prior-based variants introduce a category prior but typically follow a serial deform-then-align pipeline, where underconstrained canonical completion can corrupt correspondences and induce error cascades in pose. We propose PriorPose, a reference-guided correspondence framework that keeps the category prior explicit and solves canonicalization and alignment jointly in a shared feature space. A reference-guided seeded transformer embeds the partial observation and the category prior as token sets and fuses them via geometry-aware seeds, from which the network jointly predicts a per-point NOCS field for visible points and a canonical deformation of the prior that reconstructs a full canonical instance, while a deep pose head regresses (R,t,s) from the induced correspondences. A two-part shape consistency objective, with canonical-space and camera-space consistency losses, couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding deform-then-align error cascades. Experiments on standard and larger-category benchmarks demonstrate that PriorPose sets new state-of-the-art results on most evaluated metrics, especially under strict pose thresholds, while remaining competitive on relaxed pose and IoU metrics and showing improved robustness under shape variation and domain shift.
Yihan Chen, Huan Ren, Wenfei Yang +3
School of Information Science and Technology / National Key Laboratory of Deep Space Exploration, University of Science and Technology of China · Beijing Institute of Control Engineering
Category-level 6D object pose estimation is typically formulated as a multi-category joint learning problem with fully shared model parameters. However, pronounced geometric heterogeneity across categories entangles incompatible optimization signals in shared modules, resulting in gradient conflicts and negative transfer during training. To address this challenge, we first introduce gradient-based diagnostics to quantify module-level cross-category contention. Building on results of diagnostics, we propose DecomPose, a difficulty-aware decomposition framework that mitigates optimization contention via: (1) difficulty-aware gradient decoupling, which groups categories using a data-driven difficulty proxy and routes each instance to a group-specific correspondence branch to isolate incompatible updates; and (2) stability-driven asymmetric branching, which assigns higher-capacity branches to structurally simple categories as stable optimization anchors while constraining complex categories with lightweight branches to suppress noisy updates and alleviate negative transfer. Extensive experiments on REAL275, CAMERA25, and HouseCat6D demonstrate that DecomPose effectively reduces cross-category optimization contention and delivers superior pose estimation performance across multiple benchmarks.
Yifan Gao, Lu Zou, Zhangjin Huang +1
Hubei Key Laboratory of Intelligent Robot, Wuhan Institute of Technology, Wuhan, Hubei, China · University of Science and Technology of China, Hefei, Anhui, China · Peking University, Beijing, China