Perceiving structured shapes, such as human faces, from pixels is an inherently ambiguous task in real-world conditions. Yet, shape inference is largely posed as a deterministic regression task predicting fixed spatial coordinates. We find that deterministic regression is brittle when visual evidence is ambiguous or incomplete; under severe occlusions deterministic models exhibit structural collapse, predicting incoherent shapes or reverting to generic averages. To address this, we introduce Shape-Bayes, a probabilistic framework that couples uncertainty-aware visual perception with Bayesian shape reasoning. Rather than forcing point estimates, Shape-Bayes dynamically weights visual evidence against geometric priors to infer a structurally valid shape posterior. Demonstrated on human face shape regression, a rigorous testbed featuring complex non-rigid deformations and strict anatomical constraints, Shape-Bayes comprises: (1) a base model predicting noisy landmarks alongside distilled aleatoric uncertainties; (2) a lightweight Transformer encoding these observations into an adaptive prior over a PCA shape manifold; and (3) a differentiable Bayesian solver computing closed-form posteriors by balancing the noisy predictions against this prior. By guaranteeing complete structural integrity, Shape-Bayes achieves an absolute improvement of up to ~34% IDR over state-of-the-art deterministic models. Simultaneously, it yields highly calibrated uncertainty bounds and reduces relative error by up to 12.5%, establishing a new state-of-the-art for robust 2D face shape regression under severe occlusion. The project page is at https://shape-bayes.github.io.
Figures & tables
Figure 1: Robust Bayesian shape inference under severe occlusion. ( Left ) Deterministic regression exhibits structural collapse when forced to predict point estimates for obscured regions. ( Right ) Shape-Bayes probabilistically reconciles noisy visual evidence and a latent shape prior via aleatoric uncertainty, and infers a well-calibrated posterior. This not only guarantees structural integrity but enables the sampling of multiple plausible shape hypotheses under severe ambiguity.
Figure 2: Schematic overview of Shape-Bayes . Given an initial noisy shape estimate, Shape-Bayes dynamically balances point-wise observation likelihoods ( W ) against a data-driven structural prior ( Λ ) via a closed-form Bayesian update to recover a calibrated posterior distribution P(S) . Although demonstrated only on face shape regression here, Shape-Bayes is topology-agnostic and applicable to any other structured shape inference.
Figure 3: Shape-Bayes inference flow. (1) A base regression model predicts landmarks S′ and uncertainties σ′ , defining the likelihood precision W . (2) A Transformer maps these observations to an adaptive prior precision Λprior in PCA space. (3) A closed-form Bayesian update computes the shape posterior N(bμ,Σ~b) , yielding probable shapes via sampling and projection ( S=Sˉ+Pb ).
Figure 4: Distilled Uncertainty Supervision. Teacher predictions across affine-augmented views of an input are inverse-transformed; their empirical variance provides aleatoric targets for the base model.
Figure 5: Qualitative illustration of the Shape-Bayes inference pipeline. From left to right: input image, base model shape prediction and its aleatoric uncertainty, Bayesian posterior mean, and shapes sampled from the posterior.
Table 6
300W (Dynamic Occlusion)
WFLW (Dynamic Occlusion)
COFW (Dynamic Occlusion)
Method
Cfg.
IDR ↑
NME occ↓
NME all↓
FR ↓
AUC ↑
IDR ↑
NME occ↓
NME all↓
FR ↓
AUC ↑
IDR ↑
NME occ↓
NME all↓
FR ↓
AUC ↑
Base
83
9.87
7.06
11.8
36.8
90
12.27
9.08
23.3
30.9
–
–
–
–
–
LUVLi
+SB
100
8.86
6.67 / 6.17
9.9
39.5
100
11.33
8.73 / 7.97
22.0
32.3
–
–
–
–
–
Base
–
–
–
–
–
98
12.08
8.81
25.3
28.3
–
–
–
–
–
DSLPT
+SB
–
–
–
–
–
100
11.46
8.66 / 8.25
24.6
29.4
–
–
–
–
–
Base
78
10.36
7.07
14.4
36.5
84
10.49
8.24
21.3
33.0
100
9.56
6.61
10.8
37.1
Table 3: Generalization of Shape-Bayes (+SB) across base models and datasets ( NMEall : mean / minNME@100 ). Base models in gray natively output uncertainty, requiring no training.
Table 7: Ablation of Shape-Bayes on occluded 300W. Rows are ordered by architectural progression. Cell shading tracks performance from high failure rate ( red ) to low failure rate ( blue ).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Qualitative results on occluded 300W. HR-noSA ( Yang and Yeh, 2025 ) base model vs. Shape-Bayes .
Figure 10: Qualitative results on occluded COFW. HR-noSA ( Yang and Yeh, 2025 ) base model vs. Shape-Bayes .
Figure 11: Qualitative results on occluded WFLW. HR-noSA ( Yang and Yeh, 2025 ) base model vs. Shape-Bayes .
Base Model
Config.
Jaw
R-Brow
L-Brow
Nose
L-Eye
R-Eye
Mouth
ORFormer ( Chiang et al., 2025 )
Base
9.0
9.5
9.8
9.5
7.1
6.6
11.1
+SB
8.8
8.8
9.3
7.8
6.3
5.7
9.7
LUVLi ( Kumar et al., 2020 )
Base
10.7
14.1
12.9
9.0
10.2
9.3
9.3
+SB
9.9
11.3
11.1
7.4
8.0
7.7
8.8
STAR ( Zhou et al., 2023 )
Base
13.0
10.2
12.9
10.5
9.1
7.2
12.0
+SB
11.5
9.2
10.6
7.7
7.3
6.1
10.2
Appendix
Table 8: Region-wise shape prediction error in terms of NME occ ( ↓ ) on occluded 300W .
Base Model
Config.
Jaw
R-Brow
L-Brow
Nose
L-Eye
R-Eye
Mouth
HR-noSA ( Yang and Yeh, 2025 )
Base
11.8
11.9
11.4
11.3
9.6
10.4
12.9
+SB
11.5
11.0
10.9
9.9
8.9
9.3
11.8
ORFormer ( Chiang et al., 2025 )
Base
11.5
11.2
11.4
9.0
9.5
8.7
10.5
+SB
11.0
10.9
11.2
8.4
9.1
8.4
10.2
LUVLi ( Kumar et al., 2020 )
Base
11.5
12.9
13.6
11.3
11.5
11.1
12.8
+SB
11.1
12.2
12.7
10.2
10.5
10.2
11.8
Appendix
Table 9: Region-wise shape prediction error in terms of NME occ ( ↓ ) on occluded WFLW .
Base Model
Config.
R-Brow
L-Brow
R-Eye
L-Eye
Nose
Mouth
Chin
HR-noSA ( Yang and Yeh, 2025 )
Base
12.0
14.4
10.0
12.1
10.7
11.1
13.4
+SB
10.6
11.7
8.5
8.9
9.1
9.8
14.6
ORFormer ( Chiang et al., 2025 )
Base
10.1
11.6
8.0
8.8
9.3
9.9
8.7
+SB
9.0
10.4
7.5
7.9
8.6
9.5
8.1
STAR ( Zhou et al., 2023 )
Base
11.9
14.4
9.9
12.0
10.3
10.1
13.0
+SB
10.7
12.5
8.6
9.9
9.0
9.3
12.8
Appendix
Table 10: Region-wise shape prediction error in terms of NME occ ( ↓ ) on occluded COFW .
Object recognition (OR) in humans relies heavily on shape cues and the ability to recognize objects across varying 3D viewpoints. Unlike humans, deep networks often rely on non-shape cues such as texture and background, leading to vulnerabilities in generalization and robustness. To address this gap, we introduce ShapeY, a novel and principled benchmarking framework designed to evaluate shape-based recognition capability in OR systems. ShapeY comprises 68,200 grayscale images of 200 3D objects rendered from multiple viewpoints and optionally subjected to non-shape ``appearance'' changes. Using a nearest-neighbor matching task, ShapeY specifically probes the fine-grained structure of an OR system's embedding space by evaluating whether object views are clustered by 3D shape similarity across varying 3D viewpoints and other non-shape changes. ShapeY provides a suite of quantitative and qualitative performance readouts, including error rate graphs, viewpoint tuning curves, histograms of positive and negative matching scores, and grids showing ordered best matches, which together offer a comprehensive evaluation of an OR system's shape understanding capability. Testing of 321 pre-trained networks with diverse architectures reveals significant challenges in achieving robust shape-based recognition: even state-of-the-art models struggle to generalize consistently across 3D viewpoint and appearance changes, and are prone to infrequent but egregious matches of objects of obviously completely different shape. ShapeY establishes a principled framework for advancing artificial vision systems toward human-like shape recognition capabilities, emphasizing the importance of disentangled and invariant object encodings.
3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insufficient for inferring missing structure. We present DinoComplete, a deterministic and efficient shape completion framework that augments geometric reconstruction with voxel-aligned semantic priors distilled from DINO features. First, we construct multi-view DINO feature volumes aligned with ShapeNet data and train a student network to predict dense semantic features directly from incomplete shapes. These predicted features capture global structure and part-aware semantic context while remaining aligned with the underlying geometry. We then integrate these distilled features into a completion network, where geometric and semantic voxel representations are fused through voxel state-space modeling. To enable efficient long-range reasoning without sacrificing resolution, we introduce a multi-scale voxel Mamba module that refines the fused features by combining full-grid and chunk-wise sequence modeling. Experiments on unseen ShapeNet categories and ScanNet objects show that DinoComplete achieves stronger completion quality than prior deterministic and generative based completion methods while using fewer parameters, requiring lower memory, and achieving faster inference. Our results demonstrate that distilling semantic priors from visual foundation models improves generalization and robustness in 3D shape completion.
Furkan Mert Algan, Eckehard Steinbach
Chair of Media Technology · Munich Institute of Robotics and Machine Intelligence · School of Computation Information and Technology, Technical University of Munich
Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a generative framework for probabilistic joint estimation of object and part shapes, as well as their pose under occlusion and partial visibility from one or multiple RGB-D images. By leveraging compositional synthetic scene generation and strong 3D shape priors, RecGen generalizes across diverse object types and real-world environments. RecGen achieves state-of-the-art performance on complex, heavily occluded datasets, robustly handling severe occlusions, symmetric objects, object parts, and intricate geometry and texture. Despite using nearly 80% fewer training meshes than the previous state of the art SAM3D, RecGen outperforms it by 30.1% in geometric shape quality, 9.1% in texture reconstruction, and 33.9% in pose estimation.