cs.LGSep 30, 2026

PEG-Tab: Sampling-Time Record Repair and Release Control for Tabular Synthesis

Authors: Pengfei Li, QinYi Liu, Mohammad Khalil

Abstract

Pretrained tabular generators can reproduce training records even when aggregate utility remains high. When retraining is unavailable or too costly, sampling and release are the remaining intervention points. We present PEG-Tab (Post-Training Energy Guidance for Tabular Synthesis), a post-training repair and release-control framework for frozen tabular generators. For each generated row, a generator-native operator creates two alternatives. A shared calibrated score compares the three candidates, favours lower-risk records, and applies a final release check. We instantiate this interface for GReaT, CTGAN, TVAE, and TabDDPM without updating their parameters. Across five datasets and four generator families, PEG-Tab reduces mean Near Copy from 0.0780.078 to 0.0270.027 and lowers aggregate Exact Copy to zero. Relative to a 3×3\times post hoc filter, it retains higher utility in 12 of 16 transfer settings and Pareto-dominates the filter in eight. Gains are concentrated in copy and proximity-related risks.

Figures & tables

Explore similar work

CardsList
  1. Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training

    Apr 21, 2026Yunbo Long, Tejumade Afonja, Guangya Hao +2Post-TrainingNext-Token Distribution

  2. TabSCM: A practical Framework for Generating Realistic Tabular Data

    Apr 24, 2026Sven Jacob, Bardh Prenkaj, Weijia Shao +1Synthetic Tabular DataCausal Discovery Methods