Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porous structures within a single cubic discretization. However, voxel-based generation faces a plausibility-novelty trade-off: staying close to known geometries helps preserve geometric regularities, while moving away from them is necessary for novelty but may produce degenerate geometries. To address this challenge, we propose REGDIFF, a generative framework that couples voxel representation with latent space regulation and guided diffusion. REGDIFF introduces a repel-and-sink (RAS) mechanism to smooth the latent distribution of plausible geometries, and short-range repulsion (SRR) guidance to discourage generation overly close to known samples while maintaining geometric plausibility. We further contribute a voxel-based benchmark covering truss- and shell-type metamaterial geometries, together with an evaluation module for geometric plausibility, novelty, and diversity. Experiments show that REGDIFF outperforms voxel-based generative baselines, achieving +8.9% in geometric plausibility, +46.4% in novelty, and +128.6% in diversity on average across two datasets. These results suggest that REGDIFF is a strong geometry candidate generator for downstream evaluation. Our code is provided at https://github.com/wzhan24/ReGDiff.
Figures & tables
Figure 1 : Overview of metamaterials and benchmark development.
Figure 2 : An overview of the proposed framework ReGDiff . It encodes voxel geometries into latent space, applies RAS for latent regulation, and employs SRR-guided diffusion to generate novel yet geometrically plausible metamaterial geometry candidates.
Figure 3 : Illustration of the effect of RAS and its components.
Approaches
Geometric Plausibility Scores
Novelty Score
Diversity Score
Ssym↑
Sper↑
Scon↑
Mean ↑
Snov↑
Sdiv↑
MetaTruss (ours)
DiT-3D ( Mo et al. (2023) )
0.358
0.248
0.500
0.369
0.003
0.010
Y. Yang et al. ( Yang et al. (2024) )
0.800
0.585
0.494
0.626
0.163
0.158
XCube ( Ren et al. (2024) )
0.506
0.525
0.522
0.518
0.000
0.004
Trellis ( Xiang et al. (2025) )
0.081
0.063
0.133
0.092
0.000
0.001
Table 1 : Performance evaluation of different approaches.
Table 2 : Latent distribution visualization and generated samples with RAS regulation, contrastive regulation, or no regulation. Reg. denotes regulation, and Contra. denotes contrastive.
Approaches
Geometric Plausibility Scores
Novelty Score
Diversity Score
Ssym↑
Sper↑
Scon↑
Mean ↑
Snov↑
Sdiv↑
Case 1 (RAS + vanilla DDPM)
0.753
0.632
0.885
0.757
0.208
0.336
Case 2 (w/o reg + SRR Diff.)
0.873
0.801
0.295
0.656
0.014
0.011
Case 3 (full framework)
0.718
0.487
0.969
0.725
0.296
0.420
Table 3 : Ablation on RAS regulation and SRR diffusion.
Approaches
Geometric Plausibility Scores
Novelty Score
Diversity Score
Ssym↑
Sper↑
Scon↑
Mean ↑
Snov↑
Sdiv↑
Increase AE Param. Num.
0.753
0.471
0.977
0.734
0.308
0.411
Decrease AE Param. Num.
0.688
0.479
0.953
0.707
0.280
0.413
Increase diff. Param. Num.
0.705
0.474
0.936
0.705
0.330
0.444
Decrease diff. Param. Num.
0.722
0.493
0.955
0.723
0.286
0.409
Original setting
0.718
0.487
0.969
0.725
0.296
0.420
Table 4 : Ablation on model capacity. Param. Num. denotes parameter number.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Generated samples on MetaTruss with different models.
Figure 5 : Generated samples on MetaShell with different models.
Figure 6 : Data creation for MetaTruss.
Approaches
Geometric Plausibility Scores
Novelty Score
Diversity Score
Ssym↑
Sper↑
Scon↑
Mean ↑
Snov↑
Sdiv↑
Case 1 (RAS + vanilla DDPM)
0.935
0.884
0.930
0.916
0.305
0.625
Case 2 (w/o reg + SRR Diff.)
0.910
0.795
0.842
0.849
0.237
0.477
Case 3 (full framework)
0.923
0.856
0.978
0.919
0.380
0.783
Appendix
Table 5 : Ablation on RAS regulation and SRR diffusion.
Approaches
Geometric Plausibility Scores
Novelty Score
Diversity Score
Ssym↑
Sper↑
Scon↑
Mean ↑
Snov↑
Sdiv↑
Increase AE Param. Num.
0.915
0.858
0.985
0.919
0.362
0.711
Decrease AE Param. Num.
0.894
0.823
0.963
0.893
0.363
0.742
Increase diff. Param. Num.
0.920
0.810
0.953
0.894
0.389
0.781
Decrease diff. Param. Num.
0.907
0.852
0.956
0.905
0.359
0.775
Original setting
0.923
0.856
0.978
0.919
0.380
0.783
Appendix
Table 6 : Ablation on model capacity. Param. Num. denotes parameter number.
We introduce Discrete Voxel Diffusion (DVD), a discrete diffusion framework to generate, assess, and edit sparse voxels for SLat (Structured LATent) based 3D generative pipelines. Although discrete diffusion has not generally displaced continuous diffusion in image-like generation, we show that it can be an effective first-stage prior for sparse voxel scaffolds. By treating voxel occupancy as a native discrete variable, DVD avoids continuous-to-discrete thresholding and provides a simple framework for voxel generation, uncertainty estimation, and editing. Beyond quality gains, DVD provides more interpretable generation dynamics through explicit categorical modeling. Furthermore, we leverage the predictive entropy as a robust uncertainty metric to identify ambiguous voxel regions and complicated samples, facilitating tasks such as data filtering and quality assessment. Finally, we propose a lightweight fine-tuning strategy using block-structured perturbation patterns. This approach empowers the model to inpaint and edit voxels within a single sampling round, requiring negligible auxiliary computation and no additional model evaluations. Code is available at https://github.com/TeCai/DVD.
Zhengrui Xiang, Jiaqi Wu, Fupeng Sun +2
Imperial College London · Math Magic, Hitem3D · Math Magic
Designing 3D metamaterial microstructures that meet the intended functions remains a major challenge, as it typically requires domain expertise, iterative simulations, and extensive manual tuning. Existing work on inverse design that automatically generates microstructures based on desired target properties often suffers from limited design diversity and faces challenges in ensuring the physical feasibility of the generated structures. To address this issue, a property-informed diffusion-based network is proposed that enables the generation of 3D microstructures directly from textual descriptions. Unlike traditional property conditioning methods, our approach leverages rich guidance in terms of semantics and physical properties in the text input to support diverse structure synthesis. To enforce consistency between the generated structures and the target textual prompts, a dual alignment strategy is adopted, including contrastive text-structure alignment and test-time reward-guided alignment. Experimental results show that the model is capable of generating semantically meaningful and physically plausible structures across a wide range of material categories. Our approach has good potential for interactive microstructure design and opens up new directions for combining language-based interfaces with inverse material discovery. Code is available at: https://github.com/hongsong-wang/PropDiff-TMG
Bingxuan Dai, Hongsong Wang, Jie Gui
School of Cyber Science and Engineering, Southeast University, Nanjing 210096, China · School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China +2
We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, enabling the diffusion process to generalize across arbitrary geometries and texture parameterizations. This approach also avoids the view consistency issues inherent in video and multi-view diffusion models. Because texture space is two dimensional, we can reuse the strong priors of pretrained video diffusion models. We apply our method to high quality material reconstruction from posed photos captured under unknown lighting, as well as to text- and image guided material generation. Our method can scale to high resolutions (8K), 100+ input views, and neural material representations. In quantitative and qualitative evaluations we show state-of-the-art results for material generation and reconstruction.