Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.
Figures & tables
Figure 1
Figure 1 : Conditioned on reference image and per-part image regions derived from a segmentation map , Seg3DParts jointly generates all part meshes in a shared canonical space, yielding controllable decomposition and coherent multi-part assemblies.
Figure 2 : Overview of Seg3DParts. (a) Part-aware 3D asset generation pipeline. Seg3DParts adopts a two-stage framework. Stage 1 predicts part-aware sparse voxel structures from the input image and segmentation. Stage 2 generates geometry-aware latents for each part voxel, which are decoded into meshes aligned in a shared canonical space for direct assembly. (b) Part-aware DiT. The DiT used in both stages follows the same part-aware design. Global image features from frozen DINOv2 [ 28 ] are injected via cross-attention, while part-specific features and structural parameters modulate the network through AdaLN [ 29 ] . The DiT alternates single-part blocks for intra-part modeling and multi-part blocks for cross-part interaction, enabling coherent multi-part generation.
Figure 3 : Qualitative comparison of part-level 3D generation. We compare Seg3DParts with decomposition-based pipelines and part-structured generative models. GT part segmentations are provided as input to Seg3DParts and OmniPart. Different colors indicate different semantic parts.
Method
PartObjectNet
PartObjaverse-Tiny
FS@0.1 ↑
CD ↓
IoU ↓
FS@0.1 ↑
CD ↓
IoU ↓
PartCrafter
0.698
0.248
0.050
0.731
0.182
0.049
PartPacker
0.876
0.115
0.033
0.802
0.138
0.033
OmniPart
0.885
0.108
0.057
0.763
0.168
0.041
Hunyuan3D2.1 + PartField
0.860
0.120
0.031
0.793
0.144
0.028
Hunyuan3D2.1 + PartField + HoloPart
0.858
0.121
0.067
0.796
0.143
0.041
Table 1 : Global geometry (FS, CD) and part-overlap (IoU) evaluation on both PartObjectNet and PartObjaverse-Tiny.
Figure 4 : Ablation study on part-level generation. Removing multi-part attention leads to interpenetrating and incoherent parts, while replacing part-wise segmentation with a single object mask causes unstable decomposition. The full model produces well-separated, structurally coherent parts that closely match the ground truth.
Method
Global-Level
Part-Overlap
Part-Level
FS@0.1 ↑
CD ↓
IoU ↓
FS@0.1 ↑
CD ↓
IoU ↑
Full Model
0.856
0.129
0.054
0.631
0.467
0.649
w/o Multi-Part Interaction
0.838
0.138
0.103
0.531
0.534
0.434
Single Whole-Object Conditioning
0.840
0.137
0.126
0.488
0.589
0.444
Table 3 : Ablation study of multi-part latent interaction and segmentation-aware part conditioning. Metrics are reported for global geometry, part overlap, and part-level accuracy.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : Single-part and multi-part transformer block in Seg3DParts. Single-part block model each semantic part independently using intra-part self-attention, image cross-attention, and feed-forward layers with AdaLN conditioning. Multi-part block insert an additional cross-part attention layer to enable information exchange across parts while preserving explicit part identities. AdaLN modulation is applied in a part-specific manner, so each part’s conditioning affects only its own latent tokens.
Figure 2 : Generation of fully occluded parts. Given a single condition image (left), Seg3DParts generates a complete part decomposition (right) even for parts fully hidden in the input view, which receive an all-black image condition. Arrows indicate the recovered occluded parts: the circular pour-out opening on top of the juice tin (top) and the inner lid of the toolbox (bottom).
Figure 3 : Failure cases. Seg3DParts may produce degraded geometry when the input part segmentation is of low quality or when part boundaries are semantically ambiguous.
Part-level control is essential for modern 3D asset creation, where objects are frequently edited, reused, animated, or fabricated through their individual components. In many such workflows, users need only several specific components rather than a complete object decomposition. However, existing 3D generation methods produce all parts regardless of user intent, while promptable 3D segmentation methods typically output partial surfaces instead of reusable complete meshes. In addition, image-conditioned part generators further struggle to preserve hidden geometry and accurate placement without directly conditioning on the source mesh. To address these problems, we present SAM3D-Part, a prompt-driven framework for selective part generation from input 3D object meshes. Given a source mesh and a part prompt, SAM3D-Part first encodes the source geometry into compact mesh features and aligns them with the rendered image, selective mask, and point-map observations via pixel-wise channel fusion. The fused representation conditions a feed-forward generative model to produce only the queried component as a completed mesh. To place the generated part back into the source coordinate frame, SAM3D-Part predicts dense per-voxel correspondences and estimates the part transformation from distributed spatial evidence rather than a single global pose code. For sequential multi-part queries, previously generated parts are stored in a part cache and reused as contextual constraints, reducing conflicts among independently requested components. Extensive experiments and ablations demonstrate that SAM3D-Part can significantly improve source alignment, reduce conditioning cost, and enable consistent selective part generation, achieving state-of-the-art. Code and weights will be available at https://github.com/Jiahao620/sam3d-part.
Interactive 3D assets used in games and simulation are typically decomposed into specific semantic parts to support animation, physics, and scripted behaviors, yet most generative 3D models produce either monolithic meshes or arbitrary part decompositions that cannot be aligned with application-specific requirements. We present CubePart, a generative framework for open-vocabulary, part-controllable 3D mesh generation that exposes part structure as an explicit inference-time control signal. Given a global text prompt and a user-defined parts schema expressed as an open-ended list of part names, our method generates a set of meshes - one per schema element - that assemble into a coherent object while respecting the specified semantic structure. To enable this capability, we introduce a scalable data pipeline to construct a large open-vocabulary, part-labeled 3D dataset, along with a two-stage generative architecture that separates global shape synthesis from part-level decoding. We demonstrate that the resulting assets can be directly integrated into game engines and driven by animation and behavior scripts without manual post-processing. Project Page: https://cubepart.github.io/
Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level applications. We instead formulate part decomposition as a joint partitioning of the entire shape, where the predicted parts are non-overlapping and jointly recover the entire shape. Our key insight is that part decomposition should consider all desired parts jointly, rather than modeling each part independently. To this end, we develop a promptable model for 3D part decomposition from images or meshes. Users can specify desired parts through 3D point prompts for controllable decomposition. Given one point prompt per desired part, our model produces the corresponding parts as a complete partition of the entire shape. We build on a pretrained 3D generation model and first obtain a shape latent from either an input image or mesh. We then introduce a prompt encoder that maps each 3D point prompt to a part token while attending to the shape latent. To decode the desired parts, we propose a novel part decoder jointly scoring the entire shape against all part tokens in a coarse-to-fine manner, assigning every position within the shape volume to exactly one part. We perform part decomposition in this shared shape latent space, enabling a unified model for image-to-part generation, mesh-to-part generation, and part segmentation. Our method outperforms existing works on all part-quality metrics across all three tasks, and improves compatibility among parts by an order of magnitude over previous SOTA methods. Code and models will be released.