Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining
Organizations: School of Statistics and Data Science Shanghai University of Finance and Economics
Abstract
Causal tabular foundation models amortize effect estimation across synthetic mechanisms, but latent-effect supervision rewards posterior shrinkage rather than encoding the repeated-sample response needed in a fixed deployment population. We introduce fluctuation-supervised pretraining (FSP): each synthetic table is labeled by its average treatment effect plus its efficient influence-function fluctuation; deployment remains a frozen forward pass. Along the path , we prove an endpoint transition: every fixed retains label ambiguity of order , whereas full fluctuation makes the Gaussian label observable and reduces optimal finite-stratum causal label-prediction risk to order . A finite-pretraining bound combines label, network, episode-sampling, and optimization errors; its sampling defect controls fixed-mechanism bias, mean squared error, variance, Gaussian approximation, and, with variance-head accuracy, studentized coverage. Complementary lower bounds separate local ATE risk from the excess risk of generic finite-dictionary episode learning. Experiments trace the learned sampling response. Across 24 nonlinear continuous-covariate cells at trained context lengths, continuous-row FSP lowers checkpoint-mean macro RMSE by 7.0% versus S-learner and wins all 12 weak-overlap cells; validation-selected Summary FSP deploys faster per table in our warm one-thread benchmark. Under effect shift, matched Raw FSP lowers mean-checkpoint RMSE by 54.2% and teacher defect by 99.0% versus latent-effect supervision, and RMSE by 10.2% versus the released CausalPFN-S checkpoint. Known-effect semisynthesis tests coverage; two randomized-study evaluations show that lower RMSE can coexist with residual attenuation.
Figures & tables
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Assumption | Why it is needed | Observable diagnostic | Failure consequence |
|---|---|---|---|
| Consistency and exchangeability | Identify the ATE in ( 2 ) | Scientific design and sensitivity analysis; not testable from one table | The same algebra targets an observed standardized contrast, not necessarily a causal effect |
| and | Control empty cells and efficient-score moments | Stratum and arm counts; estimated propensity tails | Constants or variance diverge; uniform root- claim is unavailable |
| Bounded outcomes and outputs | Concentrate finite-episode squared losses | Outcome support; enforced output clipping | Replace by robust losses and explicit tail conditions |
| Independent synthetic episodes | Make the effective task sample size | Generator seeds and episode provenance | Dependent queries cannot be counted as independent tasks |
| Accurate simulator labels | Match the intended influence geometry | Teacher-corruption and nuisance diagnostics | Proposition L.1 adds bias and variance to the native defect |
| Architecture and optimization control | Separate representability, estimation, and achieved fit | Norm/width audit; held-out target loss; optimization gap | The corresponding terms in ( 12 ) persist |
| Pattern | Representative examples | Where inferential structure enters | Deployment object or operation |
|---|---|---|---|
| Amortized causal or interventional prediction | CausalPFN; Do-PFN; CausalFM [ Balazadeh Meresht et al., 2025 , Robertson et al., 2025 , Ma et al., 2026 ] | Synthetic causal or interventional target | Frozen forward prediction of an effect, uncertainty object, or interventional outcome |
| Per-data-set orthogonal or targeted estimation | DML; targeted regularization [ Chernozhukov et al., 2018 , Shi et al., 2019 ] | Nuisance fits plus an orthogonal score and cross-fitting, or a targeted neural objective | Construct an estimate anew on the analyzed data set |
| Downstream inferential construction | MP-OSPC; WALDO [ Melnychuk et al., 2026 , Masserano et al., 2023 ] | After nuisance posteriors, a predictor, or a posterior estimator has been learned | EIF-based one-step ATE posterior, or simulation-calibrated critical values and Neyman inversion |
| FSP | This work | Efficient table fluctuation and in synthetic supervision | Reuse native fixed heads; guarantees require stated mechanism-transfer and variance conditions, with a finite-pretraining upper bound and separate hard-family lower bounds |
| Method | Typical RMSE | Shifted RMSE | Boundary RMSE | Shifted | Shifted slope |
|---|---|---|---|---|---|
| Summary FSP | 0.0634 | 0.0627 | 0.0696 | 0.042 | 0.985 |
| Raw FSP | 0.0621 | 0.0624 | 0.0650 | 0.061 | 0.972 |
| Raw latent-effect | 0.0176 | 0.1364 | 0.0161 | 5.876 | -0.058 |
| CausalPFN-S | 0.0702 | 0.0695 | 0.0802 | 0.200 | 1.021 |
| S-learner | 0.0576 | 0.0565 | 0.0644 | 0.061 | 0.879 |
| T-learner | 0.0671 | 0.0664 | 0.0776 | 0.157 | 0.988 |
| Mechanism | FSP map | Bias | Coverage | Kolmogorov | ||
|---|---|---|---|---|---|---|
| Typical | Summary | .0098 | .095 | .906 | .952/.965 | .081/.082 |
| Typical | Raw | .0414 | .715 | .882 | .953/.972 | .352/.345 |
| Large effect | Summary | .115 | .909 | .963/.971 | .064/.071 | |
| Large effect | Raw | .399 | .864 | .993/.995 | .186/.199 | |
| Weak overlap | Summary | .0388 | 1.444 | .398 | .955/1.000 | .287/.311 |
| Weak overlap | Raw | .0828 | 3.286 | .328 | .706/.998 | .591/.514 |
| Target / setting | Method | Bias | RMSE | Coverage / inclusion |
|---|---|---|---|---|
| Fixed large effect | Latent-PFN | -0.1262 | 0.1271 | 0.506 |
| FSP-PFN | +0.0056 | 0.0645 | 0.953 | |
| Smoothed-stratified | -0.0010 | 0.0628 | 0.956 | |
| Stratified | +0.0006 | 0.0667 | 0.942 | |
| Fixed typical effect | Latent-PFN | -0.0143 | 0.0234 | 1.000 |
| FSP-PFN | +0.0062 | 0.0632 | 0.947 |
| Theory | Implemented check | Diagnostic readout |
|---|---|---|
| -phase transition | Seven targets, five seeds, identical generator and compute | Scaled own-label risk, FSP defect, response slope |
| Finite pretraining defect | Summary-backbone maps and aligned nine-method fixed-mechanism -grid | Bias, RMSE, |
| Bias, variance, CLT | Fixed-mechanism repetitions; native output quantiles | Centering, spread, Kolmogorov/QQ shape |
| Studentization | Native/oracle- intervals and Kolmogorov distances | , coverage, centering and shape |
| Pretraining lower bound | Exact Bernoulli proof family; separate Gaussian dictionary companion | Excess risk versus |
| Teacher robustness | Blended and shifted labels on common tables | Signed bias response |