Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data
Organizations: Betterdata AI · National University of Singapore · University of Illinois Urbana-Champaign · KAJIMA Technical Research Institute Singapore
Abstract
Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.
Figures & tables
| CoverType | Credit | Intrusion | Adult | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | Shape | Trend | MLE | Shape | Trend | MLE | Shape | Trend | MLE | Shape | Trend | MLE |
| TabTreeFormer | 0.991 | 0.991 | 0.919 | 0.929 | 0.942 | 0.971 | 0.866 | 0.982 | 0.565 | 0.935 | 0.972 | 0.822 |
| REaLTabFormer | 0.983 | 0.990 | 0.937 | 0.962 | 0.948 | 0.947 | 0.968 | 0.962 | 0.999 | 0.966 | 0.930 | 0.925 |
| TabDiff | 0.975 | 0.820 | 0.894 | 0.992 | 0.995 | 0.999 | 0.989 | 0.989 | 0.995 | 0.986 | 0.973 | 0.924 |
| GenScript | 0.984 | 0.784 | 0.467 | 0.930 | 0.897 | 0.894 | 0.874 | 0.748 | 0.572 | 0.936 | 0.858 | 0.828 |
| CoverType | Credit | Intrusion | Adult | |||||
|---|---|---|---|---|---|---|---|---|
| Method | Train | Sample | Train | Sample | Train | Sample | Train | Sample |
| TabTreeFormer | 2324.9 | 474.9 | 273.6 | 63.1 | 217.5 | 51.8 | 2117.0 | 43.6 |
| REaLTabFormer | 405.9 | 159.6 | 2194.5 | 533.5 | 581.0 | 249.9 | 175.1 | 37.8 |
| TabDiff | 5893.7 | 16.5 | 4934.3 | 13.1 | 5567.1 | 16.1 | 6760.0 | 8.6 |
| GenScript | 112.2 | 3.0 | 93.0 | 5.4 | 114.9 | 2.5 | 73.3 | 0.9 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Table | Domain | Rows | Cols. | Classes | Source |
| CoverType | forest cover | 50 000 | 55 | 7 | UCI Covertype (581 012 rows) |
| Credit | card transactions | 50 000 | 31 | 2 | Kaggle creditcardfraud (284 807) |
| Intrusion | network traffic | 50 000 | 42 | 20 | UCI KDD Cup 1999 |
| Adult | census income | 32 561 | 15 | 2 | UCI Adult (48 842 in full) |
| acc | building access | 100 | 6 | — | industrial partner |
| aicamera | camera detections | 138 341 | 4 | — | industrial partner |
| Method | Distinct pairs (real: 16) | Violating rows (%) | FD holds |
|---|---|---|---|
| TabTreeFormer | 16 | 61.104 | No |
| REaLTabFormer | 19 | 0.009 | No |
| TabDiff | 30 | 0.190 | No |
| GenScript | 16 | 0.000 | Yes |