Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we present Alchemy3D, a unified framework for training and evaluating versatile 3D asset editors that covers data construction, model architecture, and benchmark evaluation. Specifically, we curate Alchemy3D-1M, a large-scale 3D editing dataset containing 1.25M assets and 1.38M editing pairs across seven editing types. On this data, we train a family of generative flow models for general-purpose 3D asset editing. The model family supports image- and text-conditioned editing, few-step inference, and transfer to multi-view 3D part segmentation. We further introduce GEdit3D-Bench, a large-scale, open-world benchmark with a multi-dimensional evaluation protocol. Across existing and newly introduced benchmarks, our method outperforms prior methods on most metrics of editing fidelity, source preservation, and visual quality.
Figures & tables
Figure 1: Given a 3D asset and an editing instruction, our method edits the asset accordingly while preserving unrelated regions.
Dataset
Scale
PBR
Editing Types
Add
Remove
Replace
Appearance
Animation
Segmentation
Steer3D
100K
✗
✓
✓
✗
✓
✗
✗
Nano3D-100K (ICLR’26)
100K
✗
✓
✓
✓
✗
✗
✗
3DEditVerse (ICML’26)
116K
✗
✓
✓
✓
✗
✗
✗
PxForm (SIGGRAPH Asia’26)
102K
✗
✓
✓
✓
✓
✗
✗
Alchemy3D-1M
1.38M
✓
✓
✓
✓
✓
✓
✓
Table 1: Comparison of existing 3D asset editing datasets. Our Alchemy3D-1M incorporates more diverse editing types and is more than 10 × larger in scale than existing related datasets.
Figure 2: Statistics of Alchemy3D-1M dataset. Alchemy3D-1M includes 1.25M unique assets and 1.38 M editing pairs, across 7 different editing types.
Figure 3: Illustration of Alchemy3D model. Following Trellis.2, we design a three-stage editing pipeline that sequentially edits sparse structure (voxel occupancy), geometry (fine-grained surface), and material (PBR attributes). Each stage comprises hybrid attention blocks that denoise noisy target tokens by attending to source tokens from the source asset alongside condition tokens from image/text instructions. For clarity, we display decoded outputs from denoised tokens at each stage.
Figure 4: Qualitative comparison between Alchemy3D and baselines. We show only addition, removal, and replacement, the editing types supported by all baselines.
Figure 5: (a) Examples from Alchemy3D-Instruct. (b) Segmentation results from Alchemy3D-Segment, where the segmentation criterion (e.g., semantic or instance) is controlled by the image instruction. (c) Comparison between Alchemy3D and Alchemy3D-Flux.
View Qual.
Ref. Alignment
MLLM
Human
Time
Method
Aes. ↑
MUSIQ ↑
AICLIP↑
ATCLIP↑
AIDINO↑
SR ↑
VQ ↑
IF ↑
IP ↑
Pref.
(s) ↓
Nano3D
4.40
72.87
81.59
65.10
69.15
51.14%
70.16
53.84
72.09
5.1%
8.629
3DEditFormer
4.34
70.87
81.83
63.19
68.07
58.90%
66.31
59.79
63.39
5.3%
3.262
PartFlow
4.31
70.96
81.25
62.24
67.78
55.50%
65.27
56.20
66.60
6.5%
7.184
Alchemy3D-Turbo
4.58
72.94
82.36
68.31
72.88
88.18%
69.91
70.71
73.22
–
1.231
Alchemy3D
4.43
72.36
84.91
67.47
73.18
89.55%
72.74
71.38
76.19
70.7%
5.202
Table 2: Quantitative results on GEdit3D-Bench for addition, removal, and replacement. Per-type results are reported in Table 9 . Human preference scores do not sum to 100% because participants could also select a “tie” option when neither result was clearly better. Sampling time is measured on a single NVIDIA H200 GPU and averaged over 10 examples.
Add / Remove / Replace
Action
Style
Method
Aes.
AI
AT
SR
Aes.
AI
AT
SR
Aes.
AI
AT
SR
Nano3D
4.60
82.31
49.43
40.00%
–
–
–
–
–
–
–
–
3DEditFormer
4.55
81.78
49.24
70.97%
4.06
75.80
44.45
20.00%
–
–
–
–
PartFlow
4.51
81.59
49.32
67.86%
–
–
–
–
4.58
78.10
50.68
40.00%
Alchemy3D-Turbo
4.64
82.67
52.56
90.00%
4.33
79.96
51.46
90.00%
4.61
81.15
52.51
100.00%
Alchemy3D
4.60
83.77
52.10
83.87%
4.29
80.84
51.15
80.00%
4.64
82.40
54.41
90.00%
Table 3: Quantitative results on Eval3DEdit. Since Eval3DEdit originally only evaluates CLIP similarities, we additionally report aesthetic score for view quality and SR for MLLM rating.
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generative paradigms, 3D AI generation remains dominated by training-free editing pipelines. A central challenge of feedforward 3D editing lies in the lack of high-quality paired supervision. Editable 3D assets require simultaneous preservation of geometry, multi-view consistency, structural coherence, and localized edit controllability. Existing 3D editing datasets often rely on independently generated assets, image-mediated reconstruction or narrow edit taxonomies, leading to inaccurate localization, weak preservation, blurred edit boundaries, and limited semantic consistency. In this work, we introduce a new perspective: scalable feedforward 3D editing should be learned from semantic-part transformations. Based on this insight, we propose Pxform, a high-quality 3D editing dataset with over 100K consistent before/after editing pairs across seven edit types. Instead of treating objects as unstructured shapes, our pipeline grounds edits directly in semantic 3D parts. Built upon Pxform, we further propose PartFlow, a feedforward 3D editing network that injects source-aware latent control into pretrained 3D generative priors. PartFlow introduces mask-aware velocity preservation and render-space consistency supervision to jointly improve edit fidelity and source preservation, while requiring no 3D edit mask during inference. Extensive experiments demonstrate that high-quality semantic-part supervision substantially improves scalable 3D editing, enabling PartFlow to achieve state-of-the-art performance on both geometric and appearance editing benchmarks.
Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple yet accurate way to obtain 3D masks and struggle to achieve the desired edit while faithfully preserving the structure and appearance of non-target regions. To address these challenges, we present EditFlow3D, a training-free framework for local 3D editing. Given a source asset and an edit instruction, a VLM-driven workflow interprets the editing intent and automatically constructs a visual guidance image and a refined 3D editing mask, enabling localized editing in the native representation space of a pretrained 3D generative model. Specifically, mask-guided differential flow focuses the edit on the target region, while step-wise trajectory preservation maintains consistency between non-target regions and the source asset without directly replacing intermediate features. Since the existing Edit3D-Bench covers only a limited range of local editing categories, we further introduce EditFlow-Bench as a complementary benchmark encompassing a broader variety of structural and appearance edits, and evaluate EditFlow3D on both benchmarks. Quantitative results, qualitative comparisons, and a user study demonstrate that EditFlow3D achieves more accurate target-region editing and better preserves non-target regions than existing 3D editing methods.
Rui Nie, Chuang Wang, Haitao Zhou +4
Beihang University · Bambu Lab · Shanghai Jiao Tong University