cs.LGJun 8, 2026

BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

Authors: Al Zadid Sultan Bin HabibMd Younus AhamedPrashnna GyawaliGianfranco DorettoDonald A. Adjeroh

Organizations: 1,2,3,5Lane Department of Computer Science and Electrical Engineering West Virginia University, Morgantown, WV 26506, USA · 4Scientific Computing and Imaging Institute & Department of Biomedical Informatics The University of Utah, Salt Lake City, UT 84112, USA

Abstract

High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by nmn \ll m, where nn = number of samples, and mm = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in Rm\mathbb{R}^m ill-conditioned since nmn \ll m. We propose BSTabDiff, a block-subunit generative framework that partitions the mm observed features into MM latent blocks (MmM \ll m) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space RM\mathbb{R}^M while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.

Explore similar work

CardsList