Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.
Figures & tables
Figure 1: Unified 3D Partitioning from Point Prompts. (Top) We support 3D point prompts to specify desired parts and partition the input into closed, exhaustive, and mutually exclusive parts. (Right) Users can interactively control the partition through point prompts. (Bottom) We support image-to-part generation, mesh-to-part generation, and part segmentation within a single model.
Figure 2: Method overview. (a) Given a mesh or an image, a shape encoder first encodes it to shape latents Z ; the prompt encoder processes user prompts by attending to shape latents Z to get part tokens P(L) ; the part decoder scores query points x with the part tokens to predict their part assignments; mesh extraction converts these predictions into exclusive and exhaustive closed part meshes. (b, c) One layer of the prompt encoder and the part decoder, respectively.
Part quality
Part compatibility
Whole geometry
Inference
Method
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
pen% ↓
wt% ↑
CD ↓
F1@.05 ↑
time (s) ↓
Mesh input
CubePart ( Zhu et al., 2026a )
4.71
51.5
72.5
2.09
100.0
1.21
97.5
29.1
HoloPart ( Yang et al., 2026 )
5.29
46.6
73.9
1.01
32.6
1.54
95.0
99.1
X-Part ( Yan et al., 2026 )
4.53
52.5
73.4
3.05
90.8
1.19
97.3
147.2
Ours
2.73
57.0
84.4
0.06
100.0
1.56
94.1
24.4
Table 1: Comparison of our method with SOTA methods for 3D part generation from a mesh (top) and an image (bottom). We evaluate part decomposition quality, part compatibility, whole-shape geometry, and inference time.
Part quality
Inference
Method
mIoU ↑
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
time (s) ↓
SegviGen ( Li et al., 2026 )
20.50
10.05
22.8
40.3
317.9
Point-SAM ( Zhou et al., 2025 )
39.12
5.37
47.7
68.7
1.2
PartSAM ( Zhu et al., 2026b )
50.31
3.24
63.4
79.2
36.6
S 2 AM3D ( Su et al., 2026 )
50.43
3.16
62.3
80.5
0.3
P3-SAM ( Ma et al., 2026 )
54.66
4.33
61.0
74.8
19.1
Table 2: Comparison of our method with SOTA methods for 3D part segmentation. We evaluate part decomposition quality on face mIoU and on part-level CD and F1, and inference time.
Figure 3: Qualitative comparison of our method with SOTA methods on the three tasks. Part generation from an image (top) and from a mesh (middle), and part segmentation (bottom). The grey mesh beside each generated result marks in red the volume claimed by more than one part. Prior methods overlap, drop, split, or merge parts, while ours matches the ground truth.
Ablation
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
pen% ↓
(a) w/o prompt enc.
4.24
50.4
75.2
0.07
(b) w/o Lcons
3.11
53.1
80.8
0.06
(c) w/o refine.
2.86
54.6
83.9
0.05
(d) Indep. assign.
2.90
54.9
83.2
1.23
Ours
2.73
57.0
84.4
0.06
Table 3: Model ablations. We evaluate part generation from a mesh with the part-quality metrics.
Ablation
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
pen% ↓
(a) w/o prompt enc.
4.24
50.4
75.2
0.07
(b) w/o Lcons
3.11
53.1
80.8
0.06
(c) w/o refine.
2.86
54.6
83.9
0.05
(d) Indep. assign.
2.90
54.9
83.2
1.23
Ours
2.73
57.0
84.4
0.06
Table 3: Model ablations. We evaluate part generation from a mesh with the part-quality metrics.
Backbone
K
mIoU ↑
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
TripoSG
1
67.40
2.41
75.6
85.3
4
72.68
1.53
81.6
90.3
HY3D-2.1
1
69.80
2.10
78.3
87.0
4
74.84
1.35
84.0
91.8
Table 4: Ablation on the number of prompt points K . We evaluate part segmentation with the metrics of Table 2 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Symbol
Value
Architecture
Shape Latents (TripoSG / Hunyuan3D)
M
2048 / 4096
Channels
C
1024
Surface samples per shape
–
20,480
Prompt encoder layers
L
2
Part decoder layers
L′
3
Attention heads
–
8
Appendix
Table 5: Hyperparameters of our model. Symbols follow the notation of Sec. 3 , and – marks a value with no symbol in the text.
Figure 4: Training query sampling and prompt points. Small points are the sampled queries, colored by part; the larger markers are the prompt points, with balls and cubes marking two prompt sets for the same parts.
Figure 5: Part quality against the number of ground-truth parts, from a mesh (left) and from an image (right), with part quality metrics.
Figure 6: Segmentation quality against the number of prompt points per part K , for both backbones, with mIoU (left) and pCD (right). Both metrics improve steeply up to three points and change little beyond.
Figure 7: Segmentation quality against inference time per asset, pF1@.05 (left) and pCD (right), with our model at K∈{1,2,4} . Ours is the most accurate at close to the lowest time, and raising K improves quality at no cost in time.
Method
mIoU ↑
pCD ↓
pF1@.01 ↑
pF1@.05 ↑
Point-SAM ( Zhou et al., 2025 )
44.41
4.24
60.8
78.8
PartField ( Liu et al., 2025 )
52.55
7.25
53.0
65.7
S 2 AM3D ( Su et al., 2026 )
52.99
2.66
69.9
85.1
P3-SAM ( Ma et al., 2026 )
54.89
3.24
72.6
83.1
Ours
57.76
2.39
75.7
86.1
Appendix
Table 6: Comparison of our method with SOTA methods for 3D part segmentation on PartNeXt. We evaluate part decomposition quality on face mIoU and on part-level pCD and pF1.
Figure 8: Failure case and the effect of prompt placement. The same mesh is decomposed from one point per part placed near a contact between two parts (second column), from one point per part placed on the body of each part (third column), and from three points per part (fourth column).
Figure 9: Exploded views of our results. Three assets per task, shown assembled and with the parts moved apart: part generation from an image (left), from a mesh (middle), and part segmentation of a mesh (right).
Figure 10: Part segmentation on PartObjaverse-Tiny (top) and PartNeXt (bottom) . One point prompt per part, against the segmenters in Table 2 and the ground truth. Ours keeps small repeated parts separate, such as the spikes, keys, and buttons, where the baselines merge or fragment them.
Figure 11: Limitation on thin and open geometry. Large open surfaces and thin shells are poorly reconstructed by the pretrained VAE, and our decomposition consequently inherits these geometric errors. From left to right: input, VAE reconstruction, our decomposition, and ground truth.
Figure 12: Part generation from a mesh on further assets . Compared against the generators in Table 1 , with the volume claimed by more than one part marked in red beside each result. Every baseline interpenetrates at the joints, whereas our parts meet with negligible overlap and follow the ground-truth decomposition.
Figure 13: Part generation from a single image on further assets . Compared against the generators in Table 1 , with the volume claimed by more than one part marked in red beside each result. The baselines interpenetrate at the joints or return a fragmentary object, whereas our parts meet with negligible overlap and follow the ground-truth decomposition.
Figure 14: Controlling the decomposition. Objects are decomposed at three granularities, from fewer parts to more parts, through the specification each method takes: a part count for PartField, part names for CubePart, 2D masks on the image for OmniPart, and one point per part for ours. The top shows the same sample under every method; the bottom shows ours on two further samples.
Part-level control is essential for modern 3D asset creation, where objects are frequently edited, reused, animated, or fabricated through their individual components. In many such workflows, users need only several specific components rather than a complete object decomposition. However, existing 3D generation methods produce all parts regardless of user intent, while promptable 3D segmentation methods typically output partial surfaces instead of reusable complete meshes. In addition, image-conditioned part generators further struggle to preserve hidden geometry and accurate placement without directly conditioning on the source mesh. To address these problems, we present SAM3D-Part, a prompt-driven framework for selective part generation from input 3D object meshes. Given a source mesh and a part prompt, SAM3D-Part first encodes the source geometry into compact mesh features and aligns them with the rendered image, selective mask, and point-map observations via pixel-wise channel fusion. The fused representation conditions a feed-forward generative model to produce only the queried component as a completed mesh. To place the generated part back into the source coordinate frame, SAM3D-Part predicts dense per-voxel correspondences and estimates the part transformation from distributed spatial evidence rather than a single global pose code. For sequential multi-part queries, previously generated parts are stored in a part cache and reused as contextual constraints, reducing conflicts among independently requested components. Extensive experiments and ablations demonstrate that SAM3D-Part can significantly improve source alignment, reduce conditioning cost, and enable consistent selective part generation, achieving state-of-the-art. Code and weights will be available at https://github.com/Jiahao620/sam3d-part.
Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.
Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part decompositions that cannot be aligned with application-specific requirements. We present CubePart, a generative framework for open-vocabulary, part-controllable 3D mesh generation that exposes part structure as an explicit inference-time control signal. Given a global text prompt and a user-defined parts schema expressed as an open-ended list of part names, our method generates a set of meshes - one per schema element - that assemble into a coherent object while respecting the specified semantic structure. To enable this capability, we introduce a scalable data pipeline to construct a large open-vocabulary, part-labeled 3D dataset, along with a two-stage generative architecture that separates global shape synthesis from part-level decoding. We demonstrate that the resulting assets can be directly integrated into game engines and driven by animation and behavior scripts without manual post-processing. Project Page: https://cubepart.github.io/