cs.LGSep 30, 2026

MIND: Marginal-Invariant Neural Dependency Diffusion for Mixed-Type Tabular Generation

Authors: Pengfei Li, Mohammad Khalil

Organizations: Centre for the Science of Learning & Technology (SLATE), University of Bergen, Bergen, Norway

Abstract

This paper proposes MIND, a marginal-invariant neural dependency diffusion model for mixed-type tabular data. MIND does not directly learn the joint distribution in the original heterogeneous feature space. Instead, it first maps different variable types into a unified latent dependency space via column-wise marginal transport. A conditional diffusion model then learns cross-column relationships. Copula-tangent denoising separates known marginal components from learnable dependency residuals. Rank projection during the sampling phase further mitigates marginal shift in reverse diffusion. Experiments across nine diverse tabular benchmarks show that MIND consistently improves marginal fidelity and dependency preservation over existing unified approaches. By explicitly isolating marginal modelling from dependency learning, MIND achieves a strong and stable balance among marginal fidelity, joint dependency preservation, and downstream prediction utility. This work supports separating marginal and dependency modelling as a principled and highly effective paradigm for complex mixed-type tabular generation.

Figures & tables

Explore similar work

Jul 20, 2026cs.LG

Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models

Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency, which carries much of the minority-class signal in domains such as fraud detection and clinical risk. Yet standard certification is largely blind to it: a fully-factorized baseline that destroys all inter-column dependency still appears nearly real under the commonly reported linear classifier two-sample test (C2ST), and is only mildly penalized by pairwise Trend scores. This is a known weakness of linear detection scores, which we confirm on four benchmarks. We therefore decompose a stronger, gradient-boosted C2ST score into marginal, dependency, and numerical-categorical cross terms, each read against a zero-dependency reference and a real-data oracle. Applied to representative flow-matching (TabbyFlow/EF-VFM) and diffusion (TabDiff) generators, it finds a persistent dependency gap of comparable magnitude in both, tracking what their objectives share rather than anything specific to one. Dependency is necessary for minority-class utility, since a zero-dependency reference collapses it, yet the generators' residual gaps coincide with much smaller shortfalls that do not track the measured gap. The gap is neither a structural limitation of mean-field objectives nor closed by a 16x capacity increase where training is clean, which motivates supervising dependency directly in the objective as the next intervention to test.
Sep 30, 2026cs.LG

CDMD: A Cross-Dataset Mixed-Type Diffusion Model for Tabular Data

Generative models for tabular data are typically trained separately for each dataset, limiting knowledge transfer and requiring the storage of many specialized models. In this paper, we introduce CDMD, a tabular diffusion model trained jointly across heterogeneous datasets with different schemas and variable numbers of numerical and categorical features. Unlike existing cross-dataset tabular diffusion models that operate in continuous representation spaces, CDMD defines diffusion directly over the mixed-type feature space and is trained end-to-end. To accommodate heterogeneous categorical domains, we introduce a schema-restricted reverse-process parameterization for masked diffusion models, in which the output space dynamically adapts to each feature's vocabulary. We then compose numerical and categorical feature-level diffusion processes into a schema-dependent row-level process. A shared schema-aware Transformer denoiser captures dependencies between features and parameterizes the reverse process across varying schemas. On seven real-world datasets, a single jointly trained CDMD achieves the highest average generation quality among strong single-dataset and cross-dataset baselines, while using substantially fewer total parameters than the collection of separately trained models. Furthermore, pre-training on a corpus of 337 datasets improves generation on previously unseen datasets under both limited target data and limited adaptation epochs. These results demonstrate the potential of direct mixed-type diffusion for shared and transferable tabular data generation. Our code is available at https://github.com/ketatam/cdmd.
Aug 7, 2026cs.LG

FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching

Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Variational flow matching handles distinct endpoints via factorized distributions, yet leaves feature-specific processing and cross-column interactions implicit within a shared backbone. We introduce Feature-wise Unified Specialization with cross-column Exchange (FUSE) to explicitly separate these roles. FUSE applies separate adaptive mixture modules to numerical and categorical features, allowing each feature to combine shared specialized subnetworks, while joint attention preserves information exchange across all columns. We also characterize the excess population risk from restricted conditioning contexts and bound the continuous Wasserstein generation error by endpoint-prediction risk. Comprehensive experiments on eight tabular datasets demonstrate that FUSE achieves strong and consistent performance across distributional fidelity and downstream utility metrics.