This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.
Figures & tables
Figure 1: Overview of MIND: type-aware marginal mapping, conditional dependency diffusion, and calibrated inverse generation.
Metric
Indep.
G-Copula
CTGAN
TVAE
TabSyn
TabDiff
MIND
Sig. vs.
TSTR score ↑
0.344 ± 0.293
0.620 ± 0.322
0.586 ± 0.317
0.647 ± 0.345
0.576 ± 0.545
0.634 ± 0.406
0.693 ± 0.335
I
KS ↓
0.008 ± 0.005
0.008 ± 0.005
0.147 ± 0.059
0.118 ± 0.039
0.033 ± 0.022
0.028 ± 0.028
0.003 ± 0.005
C,V,S,D
JS ↓
0.012 ± 0.014
0.012 ± 0.014
0.080 ± 0.029
0.119 ± 0.080
0.027 ± 0.017
0.019 ± 0.013
0.005 ± 0.012
C,V,S,D
Column JS ↓
0.015 ± 0.011
0.017 ± 0.011
0.101 ± 0.035
0.118 ± 0.050
0.050 ± 0.033
0.043 ± 0.039
0.009 ± 0.009
C,V,S,D
Pearson err. ↓
0.104 ± 0.046
0.037 ± 0.013
0.064 ± 0.023
0.059 ± 0.029
0.017 ± 0.005
0.014 ± 0.006
0.016 ± 0.005
I,G,C,V
Pairwise MI err. ↓
0.067 ± 0.058
0.045 ± 0.040
0.033 ± 0.018
0.033 ± 0.022
0.010 ± 0.007
0.009 ± 0.009
0.008 ± 0.005
I,G,C,V
Table 1: Aggregate generation quality across nine datasets. Bold denotes the best mean in each row. The “Sig. vs.” column lists baselines significantly outperformed by MIND under two-sided paired Wilcoxon signed-rank tests on nine dataset-level seed means, with Holm correction over six comparisons within each metric ( padj<0.05 ). I denotes Indep., G denotes G-Copula, C denotes CTGAN, V denotes TVAE, S denotes TabSyn, and D denotes TabDiff.
Dataset
Indep.
G-Copula
CTGAN
TVAE
TabSyn
TabDiff
MIND
Adult
0.500 ± 0.030
0.792 ± 0.012
0.886 ± 0.001
0.884 ± 0.005
0.905 ± 0.004
0.910 ± 0.004
0.899 ± 0.003
Beijing ∗
-0.001 ± 0.003
0.233 ± 0.031
0.166 ± 0.010
0.186 ± 0.108
0.529 ± 0.022
0.560 ± 0.033
0.525 ± 0.028
Covertype
0.502 ± 0.047
0.721 ± 0.013
0.625 ± 0.069
0.882 ± 0.010
0.565 ± 0.010
0.594 ± 0.004
0.940 ± 0.001
Default
0.471 ± 0.008
0.693 ± 0.007
0.721 ± 0.016
0.734 ± 0.011
0.753 ± 0.010
0.750 ± 0.009
0.756 ± 0.008
FICO
0.529 ± 0.016
0.779 ± 0.004
0.638 ± 0.056
0.785 ± 0.006
0.784 ± 0.007
0.785 ± 0.002
0.788 ± 0.004
HeavyTail
0.452 ± 0.064
0.718 ± 0.015
0.607 ± 0.051
0.726 ± 0.015
0.727 ± 0.002
0.737 ± 0.005
0.723 ± 0.021
Table 2: Per-dataset utility under the train-on-synthetic, test-on-real protocol. Scores are averaged over XGBoost, LightGBM, and MLP, using AUC for classification datasets and R2 for regression datasets (marked with ∗ ). Entries report mean ± standard deviation over three seeds. The Avg. row averages dataset-level means. Higher is better, and bold denotes the best mean in each row.
Figure 2: Pairwise Pearson correlation errors across datasets. Lighter colours indicate more accurate dependency preservation.
Figure 3: Single-column density overlays for representative continuous variables (seed 42). The real distribution is shown in black. Selected neural baselines are compared with MIND, and each x-axis is clipped to the range between the real-data 1st and 99th percentiles for readability.
Variant
TSTR ↑
Column JS ↓
Pairwise MI error ↓
C2ST gap ↓
β -Recall ↑
Full MIND
0.844 ± 0.085
0.0135 ± 0.0109
0.0082 ± 0.0041
0.140 ± 0.104
0.9801 ± 0.0024
w/o CTD
0.843 ± 0.087
0.0136 ± 0.0109
0.0090 ± 0.0038
0.144 ± 0.103
0.9791 ± 0.0023
w/o Corr.
0.841 ± 0.089
0.0136 ± 0.0109
0.0089 ± 0.0038
0.144 ± 0.104
0.9797 ± 0.0023
w/o Attn. Extras
0.845 ± 0.082
0.0136 ± 0.0109
0.0089 ± 0.0038
0.144 ± 0.100
0.9795 ± 0.0032
w/o Projection
0.844 ± 0.089
0.0432 ± 0.0084
0.0091 ± 0.0031
0.146 ± 0.115
0.9813 ± 0.0020
Table 3: Component ablation of MIND over three datasets and three seeds. Entries report mean ± standard deviation over nine runs. Higher is better for TSTR and β -Recall, while lower is better for the remaining metrics. Bold denotes the best mean in each column.
Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency, which carries much of the minority-class signal in domains such as fraud detection and clinical risk. Yet standard certification is largely blind to it: a fully-factorized baseline that destroys all inter-column dependency still appears nearly real under the commonly reported linear classifier two-sample test (C2ST), and is only mildly penalized by pairwise Trend scores. This is a known weakness of linear detection scores, which we confirm on four benchmarks. We therefore decompose a stronger, gradient-boosted C2ST score into marginal, dependency, and numerical-categorical cross terms, each read against a zero-dependency reference and a real-data oracle. Applied to representative flow-matching (TabbyFlow/EF-VFM) and diffusion (TabDiff) generators, it finds a persistent dependency gap of comparable magnitude in both, tracking what their objectives share rather than anything specific to one. Dependency is necessary for minority-class utility, since a zero-dependency reference collapses it, yet the generators' residual gaps coincide with much smaller shortfalls that do not track the measured gap. The gap is neither a structural limitation of mean-field objectives nor closed by a 16x capacity increase where training is clean, which motivates supervising dependency directly in the objective as the next intervention to test.
Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.
Mohamed Amine Ketata, Maximilian Schambach, Stephan Günnemann
Munich Data Science Institute, Technical University of Munich, Germany · SAP SE, Germany
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. We introduce Feature-wise Unified Specialization with cross-column Exchange (FUSE) to explicitly separate these roles. FUSE applies separate adaptive mixture modules to numerical and categorical features, allowing each feature to combine shared specialized subnetworks, while joint attention preserves information exchange across all columns. We also characterize the excess population risk from restricted conditioning contexts and bound the continuous Wasserstein generation error by endpoint-prediction risk. Comprehensive experiments on eight tabular datasets demonstrate that FUSE achieves strong and consistent performance across distributional fidelity and downstream utility metrics.
Suman Cha, Seongchan Lee, Dohyun Ko +1
Department of Statistics and Data Science, Yonsei University · Department of Mathematical Sciences, KAIST