Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.
Figures & tables
Table 1: In-domain cross-dataset generation results (Section 4.1 ). Entries report mean ± standard deviation over three runs. Average is the mean over datasets. Train Set is the real-data reference. Bold and underlined entries mark the best and second-best scores in each column, excluding Train Set . GReaT cannot be applied to News because of its maximum length limit.
Figure 1: Sample-constrained generation results: Overall Evaluation Score ( S ) vs. the number of target training rows. Lines and shaded regions show the mean ± standard deviation over five runs using independently-sampled target training subsets. The final panel shows the mean across datasets.
Figure 2: Epoch-constrained generation results: Overall Evaluation Score ( S ) vs. training epochs. Lines and shaded regions indicate the mean ± standard deviation over three runs.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Overview of CDMD . Numerical features are corrupted with Gaussian noise, while categorical features are randomly masked. The tokenizer combines each noisy feature with schema information, and a shared transformer encoder captures dependencies across columns. A shared MLP predicts clean numerical values, while a shared transformer decoder uses category embeddings and the contextualized feature token to produce probabilities over each categorical column’s valid vocabulary. These predictions guide the mixed-type reverse diffusion process. The schema determines the number of feature tokens and categorical outputs, while parameters are shared across datasets and columns of the same type, allowing a single model to handle heterogeneous tables.
Dataset
# Train Rows
# Test Rows
# Cols
# Num
# Cat
# Max Cat
Task
Adult
24,421
24,421
15
6
9
42
Binary
Default
15,000
15,000
24
14
10
11
Binary
Magic
9,510
9,510
11
10
1
2
Binary
Shoppers
6,165
6,165
18
10
8
20
Binary
Diabetes
50,883
50,883
37
8
29
66
Multiclass
Beijing
20,878
20,879
8
7
1
4
Regression
Appendix
Table 2: Evaluation dataset statistics. Numerical and categorical column counts include the target. Max Cat denotes the cardinality of the largest categorical vocabulary; Binary denotes binary classification.
Quantity
Mean
Median
Min
Max
Sum
# Rows
28,585
10,999
10
100,000
9,633,139
# Cols
11.1
10
2
33
3,734
# Num
7.3
6
0
28
2,458
# Cat
3.8
2
0
23
1,276
Total datasets: 337
Appendix
Table 3: Pretraining dataset summary across the 337 datasets.
Figure 4: Distribution of row counts and total, numerical, and categorical column counts across the 337 pre-training datasets. Each dataset contributes one observation, and column counts reflect retained features after preprocessing. Row counts use 50 logarithmically spaced bins; column counts use one bin per integer.
Table 4: In-domain cross-dataset generation results for the three evaluation axes
Table 5: In-domain cross-dataset generation results for the four individual metrics of Fidelity.
Table 6: In-domain cross-dataset generation results for Privacy.
Figure 5: Few-shot performance across seven datasets and their cross-dataset average. Curves show mean performance, and shaded regions indicate one standard deviation across seeds.
Figure 6: Few-epoch performance across seven datasets and their cross-dataset average. Curves show mean performance, and shaded regions indicate one standard deviation across seeds.
This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.
Pengfei Li, Mohammad Khalil
Centre for the Science of Learning & Technology (SLATE), University of Bergen, Bergen, Norway
Tabular synthesis is critical for privacy-preserving sharing and augmentation, yet diffusion models rely on implicit mechanisms to capture inter-column relationships. We introduce Geometry-Aware Tabular Diffusion (GATD), which augments tabular diffusion denoisers with pairwise angles and lengths computed from column value differences and used as inputs and auxiliary targets. Our MLP instantiation achieves state-of-the-art benchmark performance while using 3.5x fewer parameters on average (up to 25x for classification tasks): on ten datasets, it wins 8/10 Shape, 7/10 Trend, and 9/10 downstream utility (F1/RMSE), reducing Shape and Trend error by 27% and 20%. Default loss weights transfer to GNN and Transformer denoisers, improving Shape on 27/30 and Trend on 25/30 architecture-dataset cells. A matched ablation shows supervision (not extra inputs or capacity) drives the gain. This shows explicit relational supervision is a portable inductive bias for tabular diffusion.
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. We introduce Feature-wise Unified Specialization with cross-column Exchange (FUSE) to explicitly separate these roles. FUSE applies separate adaptive mixture modules to numerical and categorical features, allowing each feature to combine shared specialized subnetworks, while joint attention preserves information exchange across all columns. We also characterize the excess population risk from restricted conditioning contexts and bound the continuous Wasserstein generation error by endpoint-prediction risk. Comprehensive experiments on eight tabular datasets demonstrate that FUSE achieves strong and consistent performance across distributional fidelity and downstream utility metrics.
Suman Cha, Seongchan Lee, Dohyun Ko +1
Department of Statistics and Data Science, Yonsei University · Department of Mathematical Sciences, KAIST