Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.
Figures & tables
Figure 1: Three families of test-time compute modify different components of the inference pipeline. Top: an inference pipeline comprising preprocessing, a frozen tabular foundation model, and combination of predictions across data views. Bottom: context construction changes the conditioning data; adaptation changes model parameters; aggregation changes the prediction pool and the rule used to select or combine its members.
Setting
Updated
Params
Δ Elo
Norm. score
Full fine-tuning
all trainable weights
58.3M
+25
0.704
LoRA ( r=16 )
low-rank updates
3.51M
+15
0.684
Value–output
WV , WO
12.9M
+19
0.693
Query–key
WQ , WK
12.9M
+22
0.696
SoftScale
query-side scaling
1.03M
+26
0.703
Table 1: Selected adaptation recipes on TabPFN-3. LoRA uses the rank with the highest Elo. Each fit contains one recipe and the same 68 references; Δ Elo is relative to the published TabPFN-3 default in that fit. Parameter counts use the regression checkpoint.
Figure 2: Per-dataset error reduction for DiagScale and full fine-tuning, relative to the same frozen backbone. Each panel orders 51 datasets by full fine-tuning’s reduction; shading marks its top five. ρ is the Spearman correlation between the two methods’ error reductions.
Backbone
Backbone params
Trainable params (% of backbone)
Δ Elo DiagScale
Δ Elo Full FT
TabPFN-3
58.3M
12.7K ( 0.022% )
+26
+25
TabICL v2
28.5M
7.3K ( 0.026% )
+82
+77
TabFM
1.65B
53.8K ( 0.0033% )
+20
+28
Table 2: DiagScale and full fine-tuning across three backbones on 51 datasets. Each Elo fit contains one adapted model and 68 fixed references (67 TabArena entries plus frozen TabFM); Δ Elo is relative to the corresponding reference baseline within that fit. Parameter counts and percentages use regression checkpoints.
Figure 3: Aggregation on TabPFN-3. A : varying native views within one configuration. B : varying configuration count under three reducers, with eight native views per configuration. Errors are relative to the released eight-view default; negative values indicate improvement. Error bars and shaded bands show 95% confidence intervals. Complete curves appear in Figure 8 .
Figure 4: Row curation on frozen, single-view TabPFN-3. A : error reduction of attention-selected versus matched random rows; 95% paired t intervals over three seeds. B/C : five-seed mean log loss across source-pool sizes. Open and filled points use separate splits from 3M-row subsets and full tables, respectively. Curve labels in B also apply to C. Dashed lines show the no-retrieval loss at 4M rows; shading marks the 4.0–4.5M memory boundary on a single 80 GB GPU.
Figure 5: Performance and runtime on TabArena. (A) TabPFN-3 recipes. (B) Each backbone’s frozen model and highest-Elo evaluated recipe, connected by an arrow. Annotations show Δ Elo relative to the corresponding reference baseline within each fit. (C) Fitting and prediction time shares for TabPFN-3 recipes. Runtime includes fitting and prediction across 51 datasets (Appendix E.4 ).
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Reference (TPU/JAX)
Ours (bf16, H100)
Δ %
Binary — reported as 1− AUC (30 datasets)
Amazon_employee_access
0.1416
0.1419
+0.24
APSFailure
0.007387
0.005329
−27.85
bank-marketing
0.2309
0.2306
−0.13
Bank_Customer_Churn
0.1229
0.1233
+0.30
Bioresponse
0.1178
0.1136
−3.55
Appendix
Table 3: Single-H100 reproduction of TabFM. Per-dataset error for the TPU/JAX reference against our bfloat16 H100 run on all 51 TabArena datasets, with Δ=(ours−reference)/reference , so a negative value means our run has the lower error. All quantities are errors, lower is better.
Dataset
OpenML ID
Total rows
Predictors
Task
Largest pool
Higgs
45570
11,000,000
28
binary
10,000,000
US_Accidents
46650
7,728,394
43
4-class
7,678,394
COMET_MC
5889
7,619,400
4
binary
7,569,400
Poker-Hand
1567
1,025,009
10
10-class
900,000
Covertype
1596
581,012
54
7-class
500,000
Appendix
Table 4: External OpenML tables used for row curation. Predictors and task type are the effective inputs seen by the model. Largest pool is the biggest context pool evaluated. The US_Accidents and COMET pools use all remaining rows after the 50,000-row holdout; Higgs is capped at 10M.
Reference entry
Imputed (%)
TabFM (frozen reproduction)
0
AutoGluon 1.5 (extreme, 4h)
0
TabPFN-3 (default)
0
TabPFN-2.6 (default)
0
TabICLv2 (default)
0
RealTabPFN-2.5 (tuned + ensembled)
0
Appendix
Table 5: The 68 reference entries used in every Elo comparison. Imputed (%) records benchmark fallback results. The published TabPFN-3 and TabICLv2 defaults are distinct from our reproduced frozen models.
Backbone
Task
Full checkpoint
DiagScale
Fraction (%)
TabPFN-3
Classification
53,153,144
13,056
0.02456
Regression
58,274,944
12,672
0.02175
TabICL v2
Classification
27,552,258
7,296
0.02648
Regression
28,544,991
7,296
0.02556
TabFM
Classification
1,639,444,298
53,760
0.00328
Regression
1,647,782,989
53,760
0.00326
Appendix
Table 6: Exact checkpoint and DiagScale parameter counts. Classification and regression use separate checkpoints. Full-model counts include fixed parameters. Percentages are rounded.
Figure 6: Learning-rate sensitivity for the cross-backbone adaptation comparison. Each point adds one candidate to the same 68 references; Δ Elo is relative to the corresponding published default. Stars mark the rates with the highest unrounded Δ Elo. All panels use five learning rates evaluated on the same 51 datasets.
Backbone
Method
Global
Five-fold (mean ± SD)
TabPFN-3
Full FT
+25.27
+21.17±2.17
TabPFN-3
DiagScale
+25.66
+19.90±2.06
TabICL v2
Full FT
+77.20
+77.20±0.00
TabICL v2
DiagScale
+82.12
+78.64±3.64
Appendix
Table 7: Learning-rate selection across datasets. Entries are Δ Elo against the published default, with 68 references and one candidate per fit. Global selects one rate using all 51 datasets; five-fold results report the mean and standard deviation across 20 dataset partitions.
Setting
Updated
Params
Δ Elo
Score
Error [95% CI]
Wins
Fallback
softscale
query scaling
1.0M
+26
0.703
−0.62%[−1.04%,−0.24%]
34/51
32%
full
all trainable weights
58.3M
+25
0.704
−0.64%[−1.07%,−0.26%]
37/51
36%
qk
query + key projections
12.9M
+22
0.696
−0.54%[−0.95%,−0.16%]
33/51
35%
middle3
middle ICL blocks 10-12
6.4M
+20
0.696
−0.60%[−1.08%,−0.19%]
36/51
34%
attn
all attention sublayers
26.2M
+20
0.693
−0.33%[−0.61%,−0.09%]
35/51
42%
mlp
all MLP sublayers
25.2M
+19
0.693
−0.37%[−0.66%,−0.11%]
38/51
42%
Appendix
Table 8: All 25 adaptation settings on TabPFN-3, each at the global learning rate with the highest Δ Elo from {10−6,10−5,10−4} . Each candidate is added separately to the same 68 references; Δ Elo is relative to the published TabPFN-3 default, and Score is TabArena’s normalized score. Error is the mean per-dataset relative change against the frozen backbone, with a 95% dataset-bootstrap interval; negative is better. Wins counts datasets with lower error. Fallback is the fraction of evaluations where validation selects pretrained weights. Parameter counts use the regression checkpoint. Rows are ordered by unrounded Δ Elo.
Backbone
Error change
95% interval
Wins
TabPFN-3
−0.091%
[−0.283%,+0.084%]
21/51
TabICL v2
−0.458%
[−1.064%,−0.016%]
26/51
TabFM
+0.102%
[−0.054%,+0.286%]
28/51
Appendix
Table 9: Paired error change of DiagScale relative to full fine-tuning; negative is better. Wins count datasets with ℓˉd<0 . The intervals describe the comparison at the selected recipes and do not establish statistical equivalence.
Backbone
Method
Median
Mean
Best
Wins
ρ (size)
TabPFN-3
full fine-tuning
−0.10%
−0.64%
−5.9%
37/51
+0.13
TabPFN-3
DiagScale
−0.08%
−0.73%
−7.8%
31/51
+0.05
TabICL v2
full fine-tuning
−0.34%
−1.30%
−10.9%
35/51
+0.26
TabICL v2
DiagScale
−0.27%
−1.60%
−11.4%
34/51
+0.26
TabFM
full fine-tuning
−0.08%
−0.32%
−5.8%
31/51
+0.16
TabFM
DiagScale
+0.01%
−0.29%
−5.8%
25/51
−0.15
Appendix
Table 10: Per-dataset relative error change against the frozen backbone, 51 datasets per row, full protocol; negative is better. Best is the largest single reduction. Wins counts datasets with any reduction. These are the deployed arms at fixed learning rates, including checkpoint selection and fallback. ρ is the Spearman correlation between training-set size and the reduction.
Figure 7: SoftScale component diagnostics. Blue circles, orange squares, and green triangles denote houses, protein, and bankruptcy. A: fraction of the joint-training gain recovered by separately training only the context-size branch s or the query-dependent gate g ; gray bars show the mean across datasets. B: each point summarizes one module group: the feature embedder or ICL stack, plus the classification decoder for bankruptcy. C: feature-embedder controls relative to its pretrained weights: trained change plus random perturbation ( Δ+R ), random perturbation alone ( R ), or reversion ( 0 ); other scale modules remain trained. Error bars in C show standard deviations over five draws for Δ+R and R . Gains in A and C are normalized by the corresponding loss reduction from jointly training both scale branches relative to pretrained weights along the same inference path.
Dataset
Train s only
Train g only
Houses
96.7%
9.2%
Protein
92.8%
16.0%
Bankruptcy
106.2%
8.1%
Appendix
Table 11: Separate branch training on the three diagnostic datasets. Entries report the percentage of the joint SoftScale training gain recovered; training both branches defines 100% . All other parameters remain frozen in each run.
Figure 8: Complete aggregation curves, including native views below the released default and random single configurations. The three reducers coincide at K=1 . Both panels use the same paired estimator and transformed 95% dataset-cluster bootstrap intervals as Figure 3 , with independent vertical ranges and all intervals visible. The additional V=48 point is included between the labeled 32- and 64-view ticks.
K
V
Reducer
Δ error [95% CI]
Wins
1
1
inner uniform
+3.74%[+0.81%,+6.70%]
4/51
1
2
inner uniform
+1.68%[+0.90%,+2.75%]
9/51
1
4
inner uniform
+0.42%[+0.05%,+0.84%]
18/51
1
8
inner uniform
— (baseline)
1
16
inner uniform
−0.48%[−0.93%,−0.14%]
33/51
1
32
inner uniform
−0.60%[−1.12%,−0.20%]
42/51
Appendix
Table 12: Every arm of Figure 8 , full protocol (816 cells), against the default baseline pbase at V=8 . Error is the percentage change 100(exp(ℓˉ)−1) over 51 equally weighted datasets with a transformed 95% dataset-cluster bootstrap interval; negative is better. Configuration rows average over independent random subsets. Wins counts datasets with reduced error and is reported for native-view settings and the complete configuration pool; dashes indicate counts not reported. Inner uniform averages native views within one configuration.
Type
Metric
Reducer
Median
Worst
Wins
binary (30)
ROC AUC
uniform
+0.30%
+5.29%
9/30
best-single
+0.05%
+4.20%
13/30
greedy
−0.31%
+1.02%
22/30
multiclass (8)
log loss
uniform
+5.80%
+164.30%
1/8
best-single
−0.85%
+0.35%
7/8
greedy
−1.21%
−0.32%
8/8
Appendix
Table 13: Relative error change against the default at K=96 , by task type; negative is better. Wins count datasets with reduced error.
Figure 9: Greedy aggregation with different fractions of training-side selector labels. All panels share the same axes and fixed member predictions. Points use the dataset-weighted mean log-error ratio against the released default; bars show transformed 95% dataset-bootstrap intervals. Pool breadth improves each curve’s point estimates, while fewer selector labels reduce the gains. Intervals exclude zero for f∈{1/2,1} and K≥16 ; all others include zero.
K
f=0.125
f=0.25
f=0.5
f=1.0
1
+11.35%
+11.35%
+11.35%
+11.35%
8
+1.37%
+0.36%
−0.35%
−0.77%
16
+0.83%
−0.19%
−0.99%
−1.45%
32
+0.61%
−0.48%
−1.34%
−1.85%
64
+0.44%
−0.72%
−1.66%
−2.27%
96
+0.39%
−0.81%
−1.77%
−2.44%
Appendix
Table 14: Greedy selection: relative error change against the default baseline pbase at V=8 , by pool size and out-of-fold label fraction; negative is better. The last row is the change from K=8 to K=96 at that budget, in percentage points. The K=1 row is the mean over randomly sampled single-member pools, not p0 ; the reducer has nothing to choose between there, so it does not vary with f .
Figure 10: Complete source-pool sweeps with single-view TabPFN-3 on five large tables, including the geometric construction omitted from the main figure. Points are the five-seed mean log losses; each panel has an independent linear vertical axis. For the three largest tables, open and filled markers distinguish the 3M-subset and full-table splits; lines connect only points from the same split. Feed-all curves contain only measured values, ending at 4M where the pool exceeds capacity. Beyond 4M, dashed lines hold that reference fixed and gray bands mark the 4.0–4.5M capacity boundary. Green lines on US_Accidents and Higgs are tuned LightGBM endpoint references (Table 17 ).
Total
Within memory
Beyond wall
Dataset
N
Eval. N
Geo.
Model-aware
Model-aware
Poker-Hand
1.03M
900k
+12.0
+87.2
—
Covertype
0.58M
500k
−67.9
+22.9
—
US_Accidents
7.73M
2.9M
−19.8
+1.5
+5.2 at 7.68M
Higgs
11M
2.9M
−4.2
+0.9
+2.2 at 6M, +3.5 at 10M
COMET_MC
7.62M
2.9M
−11.3
−2375
−2808 at 7.57M
Appendix
Table 15: Row construction on large tables, frozen TabPFN-3, single view. Relative log-loss reduction against the full-context reference for the regime, in per cent; positive is better. Within memory the reference is feed-all at the same N ; beyond the memory wall no such run exists, so it is feed-all at 4M. A dash marks a source pool that never reaches the wall. Entries summarize five seeds. The 2.9M results use the 3M-subset splits; beyond-capacity results and their 4M references use the full-table splits.
Pool
Capped union
Recursive splitting
Time ( × )
Capped
Split
6M
0.46279±0.00316
0.46266±0.00343
1.95
2.32
8M
0.46868±0.00283
0.45848±0.00328
2.20
3.16
10M
0.48530±0.00362
0.45637±0.00279
2.47
3.95
Appendix
Table 16: Higgs union-capacity comparison. Log loss is mean ± standard deviation over five seeds. Time includes attention scoring and is relative to feed-all at 4M rows. Both variants keep the retrieval budget and scoring shards fixed.
Dataset
Seed
LR
Rounds
Val. loss
Test loss
US_Accidents
0
0.02
7,972
0.14011
0.13994
1
0.02
8,093
0.14328
0.14279
2
0.02
8,056
0.13990
0.14081
3
0.02
8,411
0.13998
0.14279
4
0.02
7,479
0.14216
0.14276
Higgs
0
0.02
72,769
0.45890
0.45808
Appendix
Table 17: Validation-tuned LightGBM endpoints, five seeds. All selected models use 255 leaves. US_Accidents fits 7,524,827 rows with 153,567 validation rows; Higgs fits 10M rows with 200k additional validation rows. “Rounds” is the selected best iteration under a 100k cap.
Dataset
n
k
attn. %
union
Δ 1 view
Δ 8 views
blood-transfusion-service-center
499
4
100
499
+0.00±0.00
+0.00±0.00
diabetes
512
4
99
512
+0.00±0.85
+0.04±0.45
maternal_health_risk
676
4
91
634
+2.98±5.72
+0.77±2.43
concrete_compressive_strength
687
4
93
684
−0.57±1.96
−0.12±0.80
airfoil_self_noise
1,002
4
83
972
+0.63±3.58
−0.29±1.23
students_dropout_and_academic_success
2,950
4
50
1,814
−0.25±1.24
−0.19±0.62
Appendix
Table 18: Model-aware row construction at TabArena scale, all 15 datasets, 30 splits per dataset. n is the maximum number of training rows across splits, k the number of query clusters, attn. the mean attention mass retained by the rows selected for each query, and union the median per-cluster context size. The last two columns give mean paired relative error reductions at one view and at the released eight-view default, in per cent, with approximate 95% intervals from Equation ( 8 ); positive is better. Bold marks intervals excluding zero.
Dataset
N
Feed-all
Random ids
Attention
Attention vs. random
Covertype
500k
0.0728
0.1613
0.0558
+65.3%[+60.0,+70.6]
Poker-Hand
900k
0.3928
0.5272
0.0469
+91.1%[+89.1,+93.1]
Higgs
900k
0.4885
0.5018
0.4859
+3.2%[+2.8,+3.6]
Higgs
2.9M
0.4759
0.4911
0.4717
+3.9%[+3.5,+4.3]
Appendix
Table 19: Matched-row identity control. Log loss, mean over three seeds. The final column is the relative change from the random-identity arm to the attention arm; positive means the attention-selected rows are better. This is the contrast plotted in Figure 4 A. All four 95% intervals exclude zero.
Arm
Log loss
vs. feed-all
AUC
Rare-class share
Feed-all ( N=900 k)
0.0151
—
0.998
0.017
Attention union
0.2302
−1425%
0.950
0.403
Random identities, same size
0.0158
−4.9%
0.998
0.017
Class-marginal repair
0.1270
−741%
0.569
0.017
Appendix
Table 20: COMET_MC at N=900 k, five seeds, source pool 98.3%/1.7% . Every constructed arm refits on the same 25 clusters with a median context of 977 rows; only the contents differ. The rare-class share is the median over 125 cluster observations.
Dataset
N
∣Q∣
Draws
Union
Overlap
COMET_MC
900k
2,270
1,135,000
977
1162×
Covertype
500k
1,487
743,500
40,794
18×
Poker-Hand
900k
1,570
785,000
73,731
11×
Higgs
900k
1,383
691,500
85,188
8×
Higgs
2.9M
1,202
601,000
154,159
4×
Appendix
Table 21: Retrieval concentration within a query cluster. Query count, draws, and union size are medians over clusters and seeds (five seeds for COMET, three for the other settings). “Overlap” is the ratio of the reported median draws to median union size. Poker and Higgs statistics use the matched random-control runs, whose context sizes equal those of attention selection.
Default
Selected
F=500 controls
Dataset
Domain
D
vs. 1 view
vs. default
Per view vs. fixed
Imp. vs. random
Bioresponse
molecular
1,776
+8.9
−1.4
+4.5
+4.5
hiva_agnostic
molecular
1,617
+12.8
−4.5
+4.7
−2.7
QSAR-TID-11
molecular
1,024
+2.0
+0.5
+0.0
+2.5
CIFAR_10
image
3,072
+6.6
−2.1
+1.2
+0.0
micro-mass
mass-spec
1,300
−10.5
+35.1
−7.1
+19.2
Appendix
Table 22: Column curation on seven wide tables, frozen TabPFN-3. Entries are relative error reductions from the reference in each header; positive is better. The main selected result fixes 200 columns across all eight views for classification and 500 for QSAR-TID-11 regression, matching the default’s per-view counts. “1 view” uses all D columns. The two rightmost columns report the separate F=500 controls: fixed views each use all 500 selected columns, while per-view subsampling uses the native cap; importance and random selection are compared with 500 columns fixed across eight views. All percentages use unrounded four-seed mean errors. Task metrics and absolute errors appear in Table 23 .
No selection
Importance, F=500
Random, F=500
F=200
Dataset
Metric
1 v.
8 v.
gini
1 v.
pin
p.v.
1 v.
pin
p.v.
Imp. pin
Bioresponse
1 − AUC
0.1427
0.1300
0.1320
0.1394
0.1331
0.1270
0.1459
0.1393
0.1358
0.1319
hiva_agnostic
log loss
0.2085
0.1818
0.1840
0.2121
0.1912
0.1822
0.1951
0.1862
0.1831
0.1899
QSAR-TID-11
RMSE
0.7432
0.7282
0.7269
0.7312
0.7246
0.7244
0.7624
0.7429
0.7414
—
CIFAR_10
log loss
1.6399
1.5310
1.5531
1.6099
1.5523
1.5330
1.6267
1.5529
1.5440
1.5633
micro-mass
log loss
0.4759
0.5257
0.3358
0.4579
0.3478
0.3726
0.5579
0.4307
0.5137
0.3409
Appendix
Table 23: Complete column-curation grid: mean error over 4 seeds, in each dataset’s own metric, so values compare across a row and not down a column; lower is better. gini denotes the backbone’s importance-based per-view sampler, which also uses LightGBM gain. The Importance and Random groups both keep F=500 columns; pin gives all eight views the same set, and p.v. restores native subsampling within it. The last column adds importance selection with 200 columns fixed across eight views. Dashes denote unavailable 200-column results; regression retains the 500-column result in the main comparison.
Figure 11: Column-curation controls at a fixed 200-column selection budget with single-view TabPFN-3, separate from the eight-view comparison in Table 22 . Circles vary context rows at fixed width; squares vary width at fixed context size. The LightGBM ranking is fitted once per dataset and seed on the full training pool and held fixed through each sweep. Points are mean paired error reductions over four seeds, relative to using all available columns at that point; bars show one standard error. QSAR-TID-11 uses RMSE.
Family
Scope
Passes
Median
Example/control
Blind synthesis
51 dataset × operator
3
−4.6
none
Transduction
57 dataset × operator
14
−1.8
anneal +12.8 , airfoil +1.2
Learned rows
171 probe/dataset/dose
10
−0.1
concrete improves over placebo
Full-size joint rows
5 datasets
0
−11.8
no improvement observed
Self-stacking
19 datasets
1
−3.0
Marketing_Campaign only
OOF target encoding
9 datasets
2
−0.2
Marketing_Campaign +9.4
Appendix
Table 24: Context expansion by operator family, frozen TabPFN-3. Relative error reduction against the unmodified context, in per cent; positive is better, as in Tables 22 and 15 . The scope column names the unit being counted, which differs by family. Passes criterion counts units meeting that family’s own bar: an improvement over the unmodified context for every row except the learned-row program, where it is that program’s decision criterion.
Operator
Family
n
Median
IQR
Improves
Gaussian jitter
Blind synthesis
19
−6.19
[−11.01,−3.99]
2/19
Row interpolation
Blind synthesis
19
−2.95
[−5.49,−1.84]
1/19
Class-balanced interpolation
Blind synthesis
13
−5.90
[−20.02,−3.39]
0/13
Whole-query pseudo-labels
Transduction
19
−12.66
[−44.80,−3.44]
3/19
Cross-half pseudo-labels
Transduction
19
−2.26
[−2.76,+0.04]
5/19
Cross-half, confidence ≥0.9
Transduction
14
−0.59
[−1.41,−0.33]
2/14
Appendix
Table 25: Expansion settings and controls over the 19-dataset evaluation. Relative error reduction against the unmodified context, in per cent; positive is better. n is the number of datasets on which the operator ran; some operators apply only to specific task types or class structures. Improves counts datasets with any reduction. For learned rows, this differs from the decision criterion in Table 24 . M is the number of added rows. Nearest-row curation selects 70% of the training rows nearest each query-cluster centroid as a curation reference.
Setting
Training / construction
Deployment
Labels / selection
Blind synthesis
Jitter, pair interpolation, or class-balanced interpolation of training rows; no optimization.
D+Z , with at most n/2 added rows.
Inherited training labels; interpolated targets for regression mixup.
Transduction
Predict the query batch from D . For cross-half expansion, split queries into two parts.
D plus pseudo-labeled rows from the other half; the whole-query control adds the whole batch.
Model predictions only; optional probability or quantile-width filter.
Row controls
Duplicate training rows or sample noise anchors; no optimization.
D plus the constructed rows.
Training labels or target tails for duplication; labels for noise anchors sampled from training class proportions or fixed to the majority class.
Fold-rotated rows
Optimize free rows using real training folds plus Z as context; score a different training fold.
D+Z .
Fixed seed-row labels; training-side early stopping.
Free-row distillation
Optimize row features with Z alone as context; predict real training queries.
Same distillation objective; optimize a conditional MLP with a fixed latent bank.
D plus generated rows.
Fixed conditioning labels from training rows; training-side checkpoint selection.
Appendix
Table 26: Construction and use of expansion settings. All backbone weights are frozen. D denotes the outer-training table and Z denotes constructed rows. Only transductive settings use query features to construct the conditioning data; query labels are reserved for evaluation.
Backbone
Method
Top-5 share
Datasets improved
TabFM
DiagScale
70%
25/51
TabFM
full fine-tuning
62%
31/51
TabICL v2
DiagScale
46%
34/51
TabICL v2
full fine-tuning
48%
35/51
TabPFN-3
DiagScale
56%
31/51
TabPFN-3
aggregation
68%
41/51
Appendix
Table 27: Concentration of positive per-dataset error reductions.
Outcome
ρ32
q21
q147
q21,8
TabFM / DiagScale
0.212
0.282
0.599
0.358
TabFM / full FT
0.284
0.906
0.528
0.941
TabICL v2 / DiagScale
0.426
0.039
0.130
0.123
TabICL v2 / full FT
0.412
0.056
0.130
0.138
TabPFN-3 / DiagScale
0.339
0.317
0.444
0.264
TabPFN-3 / aggregation
0.496
0.004
0.031
0.013
Appendix
Table 28: Native-view disagreement and intervention gains. Disagreement uses TabPFN-3 predictions for all outcomes. ρ32 is its Spearman correlation with per-dataset error reduction using 32 views; q21 and q147 are Benjamini–Hochberg adjusted p -values within each outcome and across all outcomes, respectively. q21,8 repeats the within-outcome correction using eight views, four from each of the two preprocessing pipelines.
Recipe
Elo
Δ Elo
Fit (h)
Predict (h)
Total (h)
× frozen
TabPFN-3
Frozen
1658
+1
0.14
0.38
0.52
1.0
32 native views
1675
+21
0.19
1.47
1.66
3.2
DiagScale
1682
+26
6.30
0.31
6.61
12.7
Full fine-tuning
1681
+25
7.42
0.31
7.73
14.9
Aggregation (96)
1721
+67
177.71
69.79
247.50
476.4
Appendix
Table 29: Measured runtime over the same 816 TabArena cells on single H100 GPUs. Fit includes adaptation or out-of-fold member evaluation and selection; predict covers the full test batches. Pool-based methods include predictions from all 96 configurations. Elo uses 68 fixed references and one candidate per fit; Δ Elo uses the corresponding baseline within that fit. Frozen TabPFN-3 and TabICL v2 compare reproduced models with published defaults, so their deltas need not be zero. Frozen TabFM uses its rating in the TabFM full-FT fit. Runtime multiples use the reproduced frozen models.
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
We introduce TabPFN-3.5, our new flagship Tabular Foundation Model. It significantly outperforms its predecessor, TabPFN-3, and all existing baselines across a broad range of tabular problems. TabPFN-3.5 sets a new state of the art on standard tabular prediction in TabArena, and extends it to the data practitioners encounter in practice: non-i.i.d. data with temporal or grouped splits, tables with strings, text and images, high-cardinality categorical features, and wide tables with many features. These gains carry over to our task-specific harnesses: state of the art on relational data and stronger time-series forecasting. For faster inference, our variant TabPFN-3.5-Fast runs up to 3x faster than TabPFN-3 while keeping most of the accuracy gains. In addition, we upgrade TabPFN-3.5-Plus, expanding our multimodal capabilities with advanced text and date handling alongside proprietary inference optimizations. Finally, we release a new version of our Thinking mode, TabPFN-3.5-Thinking, which scales inference-time computation to push the state of the art further. It benefits from our stronger base model and from inference-time improvements that make it up to 12x faster than TabPFN-3-Thinking.
Tabular foundation models, exemplified by TabPFN, perform prediction via in-context learning, inferring test labels directly from labeled training examples. They have demonstrated competitive performance, particularly on small-to-medium datasets. However, recent tabular foundation models often improve accuracy with increasingly complex architectures, incurring higher inference cost and limiting practical deployment. In this work, we revisit the original TabPFN design and show that a lightweight row-wise attention-only backbone can remain highly competitive with two simple enhancements: a gated attention stabilization mechanism and a small set of learnable register tokens that provide global context and improve pretraining quality. The resulting model, TabSwift, supports both classification and regression, and is competitive with stronger tabular foundation models (e.g., TabPFN v2 and TabICL) while being more efficient at inference. For latency-sensitive serving, we further introduce an adaptive layer-wise early-exit mechanism that dynamically adjusts inference depth per sample. Overall, TabSwift enables efficient and anytime tabular in-context learning for practical deployments.
Si-Yang Liu, Han-Jia Ye
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China.