LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation
Authors: Hongli Xu, Zhaowei Lu, Junwen Huang, Jiaqi Hu, Peter KT Yu, Benjamin Busam, Federico Tombari, Slobodan ilic
Organizations: Technical University of Munich Munich, Germany · Munich Center for Machine Learning Munich, Germany · Siemens AG Munich, Germany · ROBOX Cambridge, MA, USA
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS correspondence, object pose and size, and a canonical Semantic Gaussian Field. Rather than treating reconstruction as a detached auxiliary task, LEGAU uses the Gaussian field as a category-conditioned structural prior that participates in multimodal feature fusion and provides global guidance for local pose reasoning. Conditioned on a categorical text embedding, LEGAU processes RGB-D observations through a transformer-based fusion module that integrates visual, geometric, and category-level cues, decoding the NOCS map, pose and size information and the Gaussian-based object representation. Extensive experiments on synthetic and real-world benchmarks show that this coupled pose-shape formulation achieves strong performance in a single-model multi-category setting, with up to 22% on SOPE and competitive transfer to real-world data. These results highlight the benefit of jointly learning canonical correspondence, object shape, and pose alignment within a unified representation.
Figures & tables
Figure 1: We present LEGAU , a unified framework for scalable category-level object pose estimation from a single view. Our method learns a Semantic Gaussian Field that provides shape- and category-aware guidance for coordinate alignment, enabling simultaneous object understanding and pose estimation, which may support downstream robotic perception tasks.
Figure 2: Pipeline of the LEGAU framework. Given an RGB-D input and a categorical text prompt, LEGAU extracts modality-specific features using DINOv2 [ 28 ] , CLIP [ 32 ] , and PointNet [ 31 ] , producing a text embedding that encodes category priors and a local embedding combining visual and geometric cues. A learnable field embedding captures global semantic and geometric priors and jointly attends with the multi-modal embeddings through a transformer. This unified representation enables coupled semantic shape understanding and pose reasoning: local embeddings drive the NOCS and pose decoders, while the global embedding guides the shape decoder to reconstruct high-level semantics and geometry as Gaussian primitives. The training of our model uses multi-level supervision, including NOCS and pose losses, photometric consistency, and a cosine-similarity objective for feature coherence.
Figure 3: Qualitative comparison between our method, AGPose [ 23 ] and GenPose++ [ 47 ] on SOPE, ROPE, and HouseCat6D. From left to right: a symmetric and occluded bowl (SOPE), a transparent glass (SOPE), a real-world toy object (ROPE), and reflective metallic tableware (HouseCat6D). These examples illustrate robustness to symmetry, occlusion, transparency, real-world novel objects, and specular surfaces.
Dataset
Method
IoU25
IoU50
IoU75
5°2cm
5°5cm
10°2cm
10°5cm
HouseCat6D [ 14 ]
VI-Net [ 6 ]
-
56.4
-
8.4
10.3
20.5
29.1
SecondPose [ 5 ]
-
66.1
-
11.0
13.4
25.3
35.7
AG-Pose [ 23 ]
88.1
76.9
53.0
21.3
22.1
51.3
54.3
LEGAU
90.7
80.6
57.4
21.4
22.7
54.0
57.4
Dataset
Method
AUC ↑ (IoU)
VUS ↑ ( n ° m cm)
IoU25
IoU50
IoU75
5°2cm
5°5cm
10°2cm
10°5cm
Table 1: Quantitative comparison on HouseCat6D [ 14 ] , SOPE [ 47 ] , and ROPE [ 47 ] . For each column, the top three methods are highlighted using a blue color map, where darker shades indicate better performance. LEGAU achieves strong performance on HouseCat6D and SOPE under the same evaluation protocol, especially in the single-model multi-category setting. On ROPE, LEGAU remains competitive in IoU-based metrics but shows a gap under stricter pose thresholds, indicating that synthetic-to-real transfer remains challenging..
Figure 4: Visual of Semantic Gaussians Reconstruction.It renders Gaussian primitives decoded in canonical object space, where the per-Gaussian semantic feature embeddings are projected to RGB colors using PCA.
GF
Sem. GF
10°5cm
IoU50
✓
✓
57.4
80.6
✓
✗
55.1 ( -4.4% )
78.3( -2.9% )
✗
✗
52.0( -9.4% )
75.0 ( -7.0% )
Table 2: Ablations on isolating the effects of Gaussian reconstruction (left) and CLIP-based category conditioning (right).
Figure 5: Qualitative comparison of shape reconstruction from partial point observations. We compare our method with a depth-based reconstruction baseline [ 45 ] and recent image-to-3D generative frameworks [ 40 , 39 , 42 ] under challenging scenarios involving heavy occlusion and transparent materials.
FoldingNet [ 41 ]
PoinTr [ 44 ]
AdaPoinTr [ 45 ]
Ours
62.72
29.87
23.17
6.72 ( -16.5% )
Table 3: Shape reconstruction results on the SOPE dataset. We evaluate on the Chamfer-L1 ( ×10−3 m, lower is better ). Best results are bolded .
School of Information Science and Technology / National Key Laboratory of Deep Space Exploration, University of Science and Technology of China · Beijing Institute of Control Engineering
Hubei Key Laboratory of Intelligent Robot, Wuhan Institute of Technology, Wuhan, Hubei, China · University of Science and Technology of China, Hefei, Anhui, China · Peking University, Beijing, China