Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information and limit the model to reproduce fine-grained details. This common design often leads to a phenomenon we term detail attenuation. Moreover, improving image-to-3D synthesis quality typically requires retraining or fine-tuning large diffusion models, which can be computationally expensive and impractical for complex 3D pipelines. In this work, we present Blended Tile Conditioning for image-to-3D generation (BTC3D), a training-free inference time framework that enhances fine-grained detail preservation in image-to-3D diffusion pipelines. To alleviate detail attenuation, we first examine the image feature additivity in image-to-3D models. Based on this property, we introduce a blended tile embedding that extracts local conditioning signals from split image regional patches, allowing the diffusion model to better preserve fine-grained visual details. To integrate the global and local conditioning guidance stably, we propose a dynamic conditioning schedule that gradually increases the influence of tile-level conditioning during later low-noise stages of diffusion. Our proposed method BTC3D operates entirely at inference time and can be seamlessly integrated into existing image-to-3D diffusion pipelines. Experimental results demonstrate that the proposed approach significantly improves texture quality and visual fidelity of the base model while maintaining global structural consistency in a training-free manner.
Figures & tables
Figure 1: Baselines vs. BTC3D . The image-to-3D models suffer from detail attenuation in 3D synthesis. The details from the input conditional images are not recognized in the 3D assets. By comparison, applying our BTC3D to TRELLIS/TRELLIS.2/Hunyuan3D-v2.1 augments the 3D assets generation by enriching such detail preservation.
Figure 2: Two key observations underpin BTC3D . (a) Compositionality: Global embedding and the average tile embedding tend to cluster in the feature space. Here, the same color denotes the same input image but with global embeddings or average tile embeddings . (b) Compatibility: Average tile embeddings improve fine-grained detail preservation for image-to-3D generation, yet it introduces minor defects in global structural consistency. For instance, the base structure of the generated hammer fails to maintain the regular square shape observed in the baseline output.
Figure 3: The overall pipeline of BTC3D , which is composed of: (a) the blended tile conditioning embedding ( BTCemb ) divides the input image into patches to enhance the feature representation; and (b) the dynamic conditioning ( DyCond ) mechanism dynamically blends global and local features aligning with the coarse-to-fine generation, addressing the detail attenuation problem.
Figure 4: Visualization of the blending weights along time steps and 3D asset generations by two different conditioning mechanisms: static conditioning vs. dynamic conditioning ( DyCond ). Compared with the static conditioning setup, our dynamic conditioning ( DyCond ) mechanism achieves further refinement of fine-grained texture details.
Method
Train Free
3D-Arena
Toys4K
PSNR ↑
SSIM ↑
CLIP-I ↑
ULIP-2 ↑
Uni3D ↑
LPIPS ↓
PSNR ↑
SSIM ↑
CLIP-I ↑
ULIP-2 ↑
Uni3D ↑
LPIPS ↓
TripoSG
✗
–
–
–
0.4091
0.3738
–
–
–
–
0.4133
0.3635
–
Hi3DGen
✗
–
–
–
0.3924
0.3653
–
–
–
–
0.4009
0.3607
–
3DTopia-XL
✗
17.0419
0.8134
0.5442
0.3319
0.3104
0.4766
16.2407
0.8218
0.5075
0.3273
0.2720
0.7928
Hunyuan3D-2.1
✗
20.6031
0.8401
0.5497
0.3773
0.3603
0.4886
20.6761
0.8351
0.6437
0.4273
0.3756
0.7805
+ BTC3D
✓
21.1871
0.8644
0.5568
0.3879
0.3719
0.4789
21.0919
0.8542
0.6535
0.4320
0.3823
0.7767
Table 1: Quantitative comparison on 3D-Arena and Toys4K benchmarks. Dashes indicate unavailable results or metrics not applicable to shape-only or untextured outputs. Higher is better for PSNR, SSIM, CLIP-I, ULIP-2, and Uni3D, while lower is better for LPIPS. Best results are shown in bold and second-best results are underlined . Note that only our method BTC3D is training-free .
Figure 5: Qualitative comparison of our method BTC3D with existing image-to-3D approaches.
Figure 6: Qualitative ablation of BTC3D . From left to right: input image, baseline, baseline with BTCemb , and BTC3D composing of BTCemb and DyCond . Red boxes highlight generation details.
Table 8
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Frequency response of Global and BTCemb . We compare full-field sinusoidal gratings with matched flat images using the normalized angular distance between their flattened complete conditioning tensors. Each point averages eight orientation–phase configurations. The shaded region highlights 128 – 448 cycles/image, and sampled frequencies are displayed at equal spacing.
Representation
Global-local Features Cos. Dist. ↑
Same-category cosine margin ↑
Average tiles
0.7466 [0.7179, 0.7758]
0.1728 [0.1433, 0.2041]
BTCemb (Ours)
0.8242 [0.8035, 0.8434]
0.2195 [0.1894, 0.2517]
Appendix
Table 6: Image feature additivity and same-category compatibility.
Figure 8: Qualitative analysis of image feature additivity. Left: PCA visualization of sampled CIELAB surface colors from the generated models, with the two input images shown as insets. Right, from top to bottom: results conditioned on the original, blended, and color-modified image embeddings.
Method
Sampling Steps
Baseline Runtime (s)
Ours Runtime (s)
Overhead (s)
Hunyuan3D-2.1
50
76.27±1.37
78.63±1.13
2.36±1.59
TRELLIS
20
2.96±0.23
4.89±0.24
1.93±0.11
TRELLIS.2
12
15.88±1.43
18.30±1.44
2.41±1.25
Appendix
Table 7: Runtime comparison of baseline pipelines before and after integrating our method on 3D-Arena ( Ebert, 2025 ) . All runtimes are measured on a single NVIDIA RTX PRO 6000 GPU.
Method
Train Free
CLIP-N BV ↑
CLIP-N ↑
CLIP-FID ↓
DINO BV ↑
DINO ↑
LPIPS BV ↓
E-Den BV ↑
E-Den ↑
Lap. Var. BV ↑
Lap. Var. ↑
3D-Arena Benchmark
TripoSG
✗
0.5725
0.4464
40.4435
–
–
–
0.0106
0.0099
98.9921
93.9380
Hi3DGen
✗
0.5693
0.4455
39.9314
–
–
–
0.0120
0.0108
99.4825
94.1028
3DTopia-XL
✗
0.5030
0.3725
25.5023
0.6326
0.3095
0.3628
0.0171
0.0159
167.3116
157.6823
Hunyuan3D-2.1
✗
0.5242
0.4034
29.3601
0.5890
0.3470
0.4102
0.0208
0.0181
149.1459
138.6304
+ BTC3D
✓
0.5263
0.4044
28.7202
0.5940
0.3539
0.4083
0.0216
0.0188
152.0670
140.6304
Appendix
Table 8: Additional semantic, distribution, perceptual, and detail metrics on 3D-Arena and Toys4K. Results are averaged over the corresponding evaluation sets. BV denotes best-view results, and E-Den denotes the Edge Density metric. Higher is better for CLIP-N, DINO, Edge Density, and Laplacian Variance, while lower is better for CLIP-FID and LPIPS. Best results are shown in bold and second-best results are underlined . For shape-only or untextured methods, DINO scores are not reported for shape-only or untextured methods.
Method
Region PSNR ↑
PSNR Win Rate
Region SSIM ↑
SSIM Win Rate
Hunyuan3D-2.1
19.6319
0.8135
+ BTC3D
20.4305
76.6% ± 4.9%
0.8309
73.5% ± 5.7%
TRELLIS
19.2718
0.8064
+ BTC3D
19.8367
69.1% ± 6.1%
0.8205
71.1% ± 5.4%
TRELLIS.2
19.8568
0.8113
+ BTC3D
20.4711
76.2% ± 4.5%
0.8470
75.1% ± 3.8%
Appendix
Table 9: Region-level paired evaluation on Toys4K. Region PSNR and SSIM are computed over spatially paired valid regions. Each merged win-rate cell reports the percentage of paired regions where BTC3D outperforms the corresponding baseline. Win rates are estimated from 10,000 sampling trials and reported as estimate ± confidence-interval half-width. Higher is better.
Method
Edge Density ↑
E-Den Win Rate
Lap. Var. ↑
Lap. Var. Win Rate
Hunyuan3D-2.1
0.0214
188.9429
+ BTC3D
0.0278
73.1% ± 4.7%
195.0055
71.9% ± 5.2%
TRELLIS
0.0159
175.3981
+ BTC3D
0.0194
67.7% ± 5.6%
180.9137
70.5% ± 4.9%
TRELLIS.2
0.0234
213.3255
+ BTC3D
0.0310
76.4% ± 4.8%
228.5734
77.1% ± 5.1%
Appendix
Table 10: Region-level paired evaluation on Toys4K. Edge Density (E-Den) and Laplacian Variance (Lap. Var.) provide auxiliary measures of edge richness and high-frequency responses in rendered images. Win rate reports the percentage of paired images where BTC3D yields a higher metric value than the corresponding baseline. Win rates are computed from 10,000 sampling trials and reported as value ± the half-width of the 95% confidence interval.
Figure 9: More qualitative comparisons by ULIP score with image-to-3D generation models.
Figure 10: Additional visualizations of BTC3D . For each example, the boxed image shows the input, the large image shows the final rendered result, and the colored image shows the corresponding normal map. The eight smaller views below provide further material and appearance analysis: the first row shows the base color, metallic, roughness, and alpha maps, while the second row presents four renderings under different realistic illumination from the same viewpoint.
Figure 11: Additional comparisons between TRELLIS.2+ BTC3D and TRELLIS.2. For each input image, the left column shows the input, while the middle and right columns present the results of TRELLIS.2+ BTC3D and TRELLIS.2, respectively. For each generated result, the left larger view shows the textured rendering, and the right larger view shows the corresponding material visualization. The four smaller views below present renderings under different realistic lighting from the same viewpoint.
Figure 12: User study interface. Desktop (left) and mobile (right) layouts of our web-based evaluation interface. The reference input image is displayed above two interactive 3D viewers labeled A and B. Participants can rotate, zoom, and pan the models, with optional synchronized camera controls, and select either “A is better” or “B is better” based on their similarity to the reference image and overall visual quality.
Figure 13: Quantitative results of the user study. (a) Overall preference distribution, where our method is preferred in 67.25% of the total 400 votes, significantly outperforming the baseline. (b) Top 10 per-item preference histogram, showing the per-sample preference rate for individual test cases. Our approach consistently receives higher user preference across diverse prompts, demonstrating the robustness and superior quality of our 3D asset generation.
N
PSNR ↑
SSIM ↑
ULIP-2 ↑
Uni3D ↑
LPIPS ↓
2
21.1799
0.8694
0.3787
0.3469
0.4615
3 (default)
21.3149
0.8721
0.3837
0.3551
0.4597
4
20.9194
0.8623
0.3805
0.3490
0.4602
5
20.7935
0.8557
0.3732
0.3378
0.4643
6
20.5157
0.8491
0.3701
0.3349
0.4710
Appendix
Table 11: Effect of the tiling parameter N on 3D-Arena. Bold values indicate the best results among the evaluated settings.
Schedule
ULIP-2 ↑
Uni3D ↑
Baseline
0.3676
0.3327
global-to-local
0.3723
0.3406
local-to-global (ours)
0.3837
0.3551
Appendix
Table 12: Effect of tile expansion direction on 3D-Arena. Higher scores are better for both metrics.
Setting
PSNR ↑
ULIP-2 ↑
Uni3D ↑
Baseline
20.6741
0.3676
0.3327
Static
20.9117
0.3733
0.3446
β =1
21.1564
0.3795
0.3488
β =2
21.3149
0.3837
0.3551
β =3
21.2213
0.3801
0.3529
Appendix
Table 13: Effect of the dynamic conditioning parameter β on 3D-Arena. Bold values indicate the best reported results.
Figure 14: Ablation of the dynamic conditioning parameter β . From left to right: input image, baseline, static tile conditioning, and results with β=1,2,3 . Red boxes highlight local details. PSNR and ULIP-2 are reported for the illustrated example.
Setting
PSNR ↑
SSIM ↑
ULIP-2 ↑
Uni3D ↑
LPIPS ↓
Baseline
20.6741
0.8480
0.3676
0.3327
0.4693
α =0.3
21.1650
0.8696
0.3801
0.3515
0.4655
α =0.4
21.3149
0.8721
0.3837
0.3551
0.4597
α =0.5
21.1094
0.8693
0.3799
0.3492
0.4649
α =0.6
20.9245
0.8601
0.3743
0.3416
0.4696
Appendix
Table 14: Effect of the fusion ratio α on 3D-Arena. Bold values indicate the best reported results.
Figure 15: Ablation of the fusion ratio α . Different fusion ratios affect local detail preservation and consistency with the input image.