Organizations: Department of Informatics, System and Communication, University of Milano-Bicocca, Milan, Italy · Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy · Department of Computer Science, University of Bonn, Bonn, Germany · Department of Information Science and Technology, Pegaso University, Naples, Italy
The decomposition of 3D point clouds into interpretable geometric primitives remains a longstanding challenge in Computer Vision and Computer Graphics. Among the available representations, superquadrics offer a compact and expressive model capable of capturing a wide range of shapes. However, their estimation is inherently challenging, as it requires solving a non-linear optimization problem and is particularly sensitive to noise, outliers, and overlapping structures. While robust estimation methods such as RANSAC and its variants achieve strong performance, they rely primarily on spatial proximity and residual-based criteria, often leading to incorrect inlier assignments across adjacent or complex arrangements of primitives. In this work, we introduce a geometric-aware framework for primitive decomposition that explicitly incorporates local surface properties into the fitting process. Specifically, we propose an inlier refinement step formulated as an energy minimization problem and solved via graph-cut optimization. Our formulation integrates geometric priors, such as normal consistency, enabling more reliable inlier selection beyond purely residual-based criteria. The approach naturally applies to both single-model estimation and multi-model decomposition. By leveraging geometric information beyond point-wise residuals, our method reduces erroneous inlier propagation and stabilizes parameter estimation. Experiments on synthetic and real datasets show consistent improvements in geometric accuracy, robustness to noise and outliers, and convergence efficiency compared to state-of-the-art RANSAC-based methods.
Figures & tables
Figure 1 : Primitive Decomposition of different 3D point clouds (top row) exploiting our Geometric Aware Inlier Refinement (GAIR) to extract superquadrics (bottom row).
Figure 2 : Geometric ambiguity in 3D primitive decomposition and inlier refinement strategies. (a) Two spatially close points may belong to different surface regions, as indicated by their differing normals in the green circle. (b) Initial inlier set Ij obtained via residual-based thresholding. (c) Proximity-based refinement propagates inliers across adjacent surfaces due to proximity. (d) Incorporating geometric consistency prevents assignments across incompatible regions, yielding a coherent inlier set.
Figure 3 : Our geometric-aware refinement in a nutshell. Given a candidate model hj and its initial inlier set Ij , GAIR produces a refined set Ij that better adheres to the underlying surface geometry.
Figure 4 : Overview of the datasets used in our experiments. From left to right: synthetic TangentSuperquadrics point clouds, used to evaluate pure fitting accuracy; SqSoup in both the single-model and multi-model configuration, used to test the algorithm in a setup with few points; a subset of CAD-like shapes from Thingi10K , showcasing the applicability of our method to more complex and realistic geometries, and, finally, a set of four 3d scans from Sketchfab . We also considered real LiDAR data reported in Fig. 11
Component
Parameter
Value
RanSaC
Inlier threshold, synthetic data
2.5σ
Inlier threshold, real scans
0.015m
Outer iterations
20
Local opt.
Inner iterations
25
Graph
Neighbors, controlled data
k=6
Neighbors, real scans
k=10
Table 1 : Default parameters used in the experiments.
Scale s
Method
CD ↓
HD ↓
1.0
Vanilla RanSaC
0.64 ± 0.10
2.15 ± 0.45
GAIR- RanSaC (ours)
0.50 ± 0.01
1.73 ± 0.28
1.5
Vanilla RanSaC
0.70 ± 0.07
2.06 ± 0.22
GAIR- RanSaC (ours)
0.43 ± 0.04
0.96 ± 0.16
2.0
Vanilla RanSaC
0.67 ± 0.07
2.01 ± 0.10
GAIR- RanSaC (ours)
0.44 ± 0.02
1.62 ± 0.21
Table 2 : Sensitivity to the inlier threshold ε=sσ on TangentSuperquadrics . Chamfer Distance (CD), Hausdorff Distance (HD) averaged over runs. Best values in bold.
Outlier
GAIR
GC
LO
RanSaC
0%
0.0260
0.0478
0.0480
0.0595
5%
0.0302
0.0569
0.0590
0.0823
10%
0.0226
0.0493
0.0522
0.1113
15%
0.0302
0.0583
0.0597
0.1143
20%
0.0318
0.0594
0.0701
0.1465
25%
0.0361
0.0700
0.0754
0.1425
Table 3 : Single-model fitting: mean IAE across the three single model point clouds from SqSoup and 25 trials per condition. Best result per row in bold .
Figure 5 : Average IAE as a function of the outlier ratio on the single-model subset of SqSoup .
Figure 6 : IAE versus execution time for the single-model subset of SqSoup .
(σ,Nout)
Method
CD ↓
HD ↓
IAE ↓
(0.1,0)
GC- RanSaCov
0.82 ± 1.07
5.09 ± 8.87
0.06 ± 0.06
GAIR- RanSaCov (ours)
0.29 ± 0.01
0.58 ± 0.03
0.03 ± 0.00
(0.1,4000)
GC- RanSaCov
0.38 ± 0.20
1.59 ± 1.91
0.08 ± 0.09
GAIR- RanSaCov (ours)
0.27 ± 0.00
0.59 ± 0.12
0.03 ± 0.00
(0.2,0)
GC- RanSaCov
0.29 ± 0.01
1.51 ± 0.49
0.09 ± 0.00
GAIR- RanSaCov (ours)
0.26 ± 0.01
0.91 ± 0.20
0.08 ± 0.00
Table 4 : Primitive decomposition on the synthetic scene tangent superquadrics consisting of κ=4 intertwined superquadrics. Best values in bold .
Figure 7 : Qualitative results of primitive decomposition using GAIR- RanSaCov on shapes from the 3D Shapes dataset. (a) Robustness under severe outlier contamination: despite strong noise, the decomposition remains consistent with the underlying geometry, as normal coherence prevents incorrect inlier assignments. (b) Example on a bird point cloud, illustrating the ability of the method to recover meaningful primitives on complex shapes.
Figure 8 : Primitive decomposition results across datasets. (a) and (b) report the average Chamfer Distance (CD) and Hausdorff Distance (HD) on SqSoup and 3D Shapes . (c) shows the Inlier Assignment Error (IAE) across both datasets. GAIR- RanSaCov consistently achieves lower error and better geometric accuracy.
Figure 9 : Qualitative comparison on publicly available Sketchfab point clouds. Within each group, from left to right, we show the input point cloud, the decomposition obtained with GC- RanSaCov , and the decomposition obtained with GAIR- RanSaCov .
Figure 10 : Ablation studies on the energy terms, tested on the Hammer point cloud.
Figure 11 : Qualitative comparison on 3D real scanned point clouds captured with an iPhone 16 Pro equipped with LiDAR. Within each group, from left to right, we show the input point cloud, the decomposition obtained with GC- RanSaC , and the decomposition obtained with GAIR- RanSaC .
Superquadrics have proven to provide a compact, geometrically meaningful representation for 3D objects. However, existing methods suffer from limited reconstruction accuracy, are restricted to rigid primitives, and lack robustness to partial point clouds. In this work, we present SuperFlex, an enhanced framework that expands the expressive power and applicability of superquadric decompositions. First, we introduce a novel loss formulation which significantly improves reconstruction accuracy. Second, we include bending and tapering deformations, enabling high-fidelity representation of curved and asymmetric geometries. Finally, we leverage these high-quality decompositions as supervision to train a model that is robust to partial real-world point clouds. Experiments demonstrate substantial improvements in reconstruction accuracy over both optimization- and learning-based baselines while maintaining a highly compact primitive representation.
Gabriel Tavernini, Elisabetta Fedele, Tiago Novello +3
This work presents a novel method for fitting superquadrics to point clouds under the contamination of noise and outliers, which has many applications for shape modeling across diverse fields. Unlike prior approaches that either exclusively focus on fitting rigid or deformable superquadrics, or suffer from robustness and numerical instability issues, our method redefines the problem from a new unsupervised clustering perspective, enabling the holistic fitting of both rigid and deformable superquadrics within a unified framework. Central to our approach is a stable optimization function inspired by unsupervised clustering analysis, where we formulate the point cloud data and samples from the potential parametric surface as clustering members and centroids, respectively. Then, the clustering process with dynamic updates to centroid locations serves as a direct proxy for optimizing superquadric parameters, establishing a principled link between geometric fitting and clustering dynamics. We further derive the relationship between pairwise computations of clustering centroids and clustering members to orthogonal distances, effectively eliminating the need for the time-consuming surface sampling process. Moreover, our formulation provides closed-form analytical solutions for both the fuzzy membership degree vector and the covariance matrix, ensuring efficient iteration optimization and enabling more effective handling of geometric deformations. In addition, we provide a theoretical certificate of convergence analysis and demonstrate that the clustering-inspired fitting method can escape local minima by inherently increasing the convexity of the objective function. The implementation is publicly available at https://github.com/zikai1/SuperquadricFitting.
Mingyang Zhao, Sipu Ruan, Xiaohong Jia
University of Chinese Academy of Sciences · Robotics Institute, School of Mechanical Engineering and Automation, Beihang University
Point clouds are a fundamental representation for robotic perception tasks such as localization, mapping, and object pose estimation. However, LiDAR-acquired point clouds are inherently sparse and non-uniform, providing incomplete observations of the underlying scene geometry. This makes reliable geometric reasoning challenging and degrades downstream perception performance. Existing approaches attempt to compensate for these limitations by estimating local geometry, but often rely on hand-crafted statistics or end-to-end supervised learning, which can suffer from limited scalability or require large amounts of accurately labeled data. To address these challenges, we explicitly model point cloud geometry under a principled mathematical formulation. We represent local geometry as a statistical manifold induced by a family of Gaussian distributions, where each point is associated with a Gaussian capturing its local geometric structure. Based on this formulation, we introduce Point-to-Ellipsoid (POLI), a deep neural estimator that predicts per-point Gaussian geometry. POLI learns a mapping from point cloud observations to their underlying geometry in a self-supervised manner, removing the need for labeled data while preserving strong geometric inductive biases. The resulting representation integrates seamlessly into existing robotic perception pipelines without architectural modifications. Extensive experiments show that POLI enables accurate and robust geometry estimation and consistently improves performance across diverse robotic perception tasks.
Jinwoo Lee, Jiwoo Kim, Woojae Shin +2
Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea · Daegu Gyeongbuk Institute of Science and Technology (DGIST), Daegu, Republic of Korea