This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.
Figures & tables
Figure 1: Overview of MIND: type-aware marginal mapping, conditional dependency diffusion, and calibrated inverse generation.
Metric
Indep.
G-Copula
CTGAN
TVAE
TabSyn
TabDiff
MIND
Sig. vs.
TSTR score ↑
0.344 ± 0.293
0.620 ± 0.322
0.586 ± 0.317
0.647 ± 0.345
0.576 ± 0.545
0.634 ± 0.406
0.693 ± 0.335
I
KS ↓
0.008 ± 0.005
0.008 ± 0.005
0.147 ± 0.059
0.118 ± 0.039
0.033 ± 0.022
0.028 ± 0.028
0.003 ± 0.005
C,V,S,D
JS ↓
0.012 ± 0.014
0.012 ± 0.014
0.080 ± 0.029
0.119 ± 0.080
0.027 ± 0.017
0.019 ± 0.013
0.005 ± 0.012
C,V,S,D
Column JS ↓
0.015 ± 0.011
0.017 ± 0.011
0.101 ± 0.035
0.118 ± 0.050
0.050 ± 0.033
0.043 ± 0.039
0.009 ± 0.009
C,V,S,D
Pearson err. ↓
0.104 ± 0.046
0.037 ± 0.013
0.064 ± 0.023
0.059 ± 0.029
0.017 ± 0.005
0.014 ± 0.006
0.016 ± 0.005
I,G,C,V
Pairwise MI err. ↓
0.067 ± 0.058
0.045 ± 0.040
0.033 ± 0.018
0.033 ± 0.022
0.010 ± 0.007
0.009 ± 0.009
0.008 ± 0.005
I,G,C,V
Table 1: Aggregate generation quality across nine datasets. Bold denotes the best mean in each row. The “Sig. vs.” column lists baselines significantly outperformed by MIND under two-sided paired Wilcoxon signed-rank tests on nine dataset-level seed means, with Holm correction over six comparisons within each metric ( padj<0.05 ). I denotes Indep., G denotes G-Copula, C denotes CTGAN, V denotes TVAE, S denotes TabSyn, and D denotes TabDiff.
Dataset
Indep.
G-Copula
CTGAN
TVAE
TabSyn
TabDiff
MIND
Adult
0.500 ± 0.030
0.792 ± 0.012
0.886 ± 0.001
0.884 ± 0.005
0.905 ± 0.004
0.910 ± 0.004
0.899 ± 0.003
Beijing ∗
-0.001 ± 0.003
0.233 ± 0.031
0.166 ± 0.010
0.186 ± 0.108
0.529 ± 0.022
0.560 ± 0.033
0.525 ± 0.028
Covertype
0.502 ± 0.047
0.721 ± 0.013
0.625 ± 0.069
0.882 ± 0.010
0.565 ± 0.010
0.594 ± 0.004
0.940 ± 0.001
Default
0.471 ± 0.008
0.693 ± 0.007
0.721 ± 0.016
0.734 ± 0.011
0.753 ± 0.010
0.750 ± 0.009
0.756 ± 0.008
FICO
0.529 ± 0.016
0.779 ± 0.004
0.638 ± 0.056
0.785 ± 0.006
0.784 ± 0.007
0.785 ± 0.002
0.788 ± 0.004
HeavyTail
0.452 ± 0.064
0.718 ± 0.015
0.607 ± 0.051
0.726 ± 0.015
0.727 ± 0.002
0.737 ± 0.005
0.723 ± 0.021
Table 2: Per-dataset utility under the train-on-synthetic, test-on-real protocol. Scores are averaged over XGBoost, LightGBM, and MLP, using AUC for classification datasets and R2 for regression datasets (marked with ∗ ). Entries report mean ± standard deviation over three seeds. The Avg. row averages dataset-level means. Higher is better, and bold denotes the best mean in each row.
Figure 2: Pairwise Pearson correlation errors across datasets. Lighter colours indicate more accurate dependency preservation.
Figure 3: Single-column density overlays for representative continuous variables (seed 42). The real distribution is shown in black. Selected neural baselines are compared with MIND, and each x-axis is clipped to the range between the real-data 1st and 99th percentiles for readability.
Variant
TSTR ↑
Column JS ↓
Pairwise MI error ↓
C2ST gap ↓
β -Recall ↑
Full MIND
0.844 ± 0.085
0.0135 ± 0.0109
0.0082 ± 0.0041
0.140 ± 0.104
0.9801 ± 0.0024
w/o CTD
0.843 ± 0.087
0.0136 ± 0.0109
0.0090 ± 0.0038
0.144 ± 0.103
0.9791 ± 0.0023
w/o Corr.
0.841 ± 0.089
0.0136 ± 0.0109
0.0089 ± 0.0038
0.144 ± 0.104
0.9797 ± 0.0023
w/o Attn. Extras
0.845 ± 0.082
0.0136 ± 0.0109
0.0089 ± 0.0038
0.144 ± 0.100
0.9795 ± 0.0032
w/o Projection
0.844 ± 0.089
0.0432 ± 0.0084
0.0091 ± 0.0031
0.146 ± 0.115
0.9813 ± 0.0020
Table 3: Component ablation of MIND over three datasets and three seeds. Entries report mean ± standard deviation over nine runs. Higher is better for TSTR and β -Recall, while lower is better for the remaining metrics. Bold denotes the best mean in each column.