Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label. However, whether this fixed-weight account describes pretrained tabular foundation models (TFMs) remains unclear. Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map. We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations. Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing. We further find that joint processing emerges with training and that attention scores carry most of the measured interaction. Together, these findings motivate TFM explanations that account for how context labels change the influence of individual examples.
Figures & tables
Figure 1: Labels are processed jointly, and attention carries most of the measured interaction. a : three ways a model could use context labels while the features and the query stay fixed. The blue arrow is the influence ai of row i on the internal prediction y^int . With fixed weights, this influence never changes. If labels are handled one at a time, a label can change only its own influence. In joint processing, another label y~j changes how much row i matters, which is what all five models show (§ 4.2 ). b : inside an attention step, labels can act through the attention scores and through the representations that those scores combine. A stop-gradient, applied at every attention step, lets the scores pass unchanged but stops labels from influencing them when derivatives are taken. The prediction stays the same, while most of the measured interaction disappears (§ 4.3 ). The stop-gradient shows which derivative paths carry the interaction in the computation as implemented, not that these paths are necessary.
Figure 2: The standardization trap and the two certificates (§ 3.2 ). a : the network only ever sees labels on the standardization manifold M , but derivatives also react to label changes that would leave it. For a label perturbation v∈Rn , the projector PT keeps the part that slides along M , so that a small step keeps the label mean and changes the label norm only negligibly. b : each shape shows how strongly the prediction curves along every direction u within the tangent space, drawn as the distance from the center. This curvature splits into a part that is equal in every direction, λPT , and the remainder T . A function such as ga that is affine on M can produce only the first part (Lemma 1 ), so AT=∥T∥F measures what such a map cannot explain. The equal part is enlarged for visibility. c : S contains every curvature pattern that arises when each label is transformed separately, as in ga,ϕ . PST is the closest such pattern to T , and the length of what is left over is Xsep (Lemma 2 ). d : the same steps on a constructed Hessian. It sums terms that act only off M , a separate effect for each label, and couplings between pairs of labels. Only the pair couplings reach T−PST .
Quantity
What it keeps
ga
ga,ϕ
Joint map
Hessian H
every label change
2qI
2qI+D
=0
Tangent curvature PTHPT
changes that stay on M
2qPT
2qPT+PTDPT
=0
T , of size AT
minus the isotropic part
0
in S , =0
=0
T−PST , of size Xsep
minus the one-label-at-a-time part
0
0
=0
Table 1: From the Hessian to the two certificates. Each row applies one more correction to the Hessian H of the map. The last three columns give the result for ga , which is a fixed-weight map on M (Eq. 2 ), for the one-label-at-a-time map ga,ϕ (Eq. 4 , with D=diag(ϕk′′(y~k)) ), and for a map in which labels interact. Only AT vanishes for every fixed-weight map, and only Xsep also vanishes for every one-label-at-a-time map.
Share (%)
Attention drop (%)
Model
AˉT
Xˉsep
AT
Xsep
TabICL-v2
98.7
63.1
81.6
82.0
TabPFN-v3
98.5
70.8
85.9
90.0
TabPFN-v2.6
97.7
73.4
90.9
95.3
TabFM
99.2
67.5
78.2
82.1
TabDPT
98.3
40.4
68.0
75.3
Table 2: All five models process labels jointly, and most of the interaction passes through attention. The shares (Eq. ( 6 )) and the relative fall of AT and Xsep when derivatives cannot pass through the attention scores (Section 4.3 ), in percent, each averaged over the four datasets. A fixed-weight map would give AˉT=0 , and a one-label-at-a-time map would give Xˉsep=0 . Appendix Table 8 gives every setting.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Control
Δperm
∥PTHPT∥F
AT
Xsep
LS
Δperm⊥
Affine on M
1.21
38.7
4×10−14
3×10−14
6×10−14
7×10−14
Separable on M
1.23
9.7
1.11
3×10−14
0.36
0.53
Coupled
0.96
0.73
0.72
0.72
0.75
0.96
Appendix
Table 3: Calibration on maps with known behavior. Values use the same estimator as the models, at n=256 in float64. The raw statistics Δperm and ∥PTHPT∥F can be nonzero even when the map is affine on M , which is the trap of Section 3.2 made numerical. Only Xsep is also zero on a map that transforms one label at a time. The first-order statistics LS and Δperm⊥ are zero on every map that is affine on M .
Model
Identifier
Readout
TabICL-v2 ( Qu et al., 2026 )
tabicl-regressor-v2
linear mean-over-quantiles
TabPFN-v3 ( Grinsztajn et al., 2026 )
ModelVersion.V3
bucket-softmax
TabPFN-v2.6 ( Grinsztajn et al., 2025 )
ModelVersion.V2_6
bucket-softmax
TabFM 1.0 ( Kong and Das, 2026 )
tabfm-1.0.0 (regression)
scalar regression output
TabDPT v1.2 ( Hosseinzadeh et al., 2026 )
tabdpt1_2.safetensors
bucket-softmax
Appendix
Table 4: The five evaluated models and the readout we differentiate. The identifier names the checkpoint or model version that we load, and the readout names the function of the network’s output that supplies the internal prediction y^int rather than the package’s default combined prediction. Every derivative computation in this paper uses this single readout, except the readout-head check of Appendix C.1 .
Experiment
Grid and settings
Replication
First-order tests
Core 8, 4 models, 32
n=256 , 3 contexts, nq=8 , 6 permutations
Projected Hessian
Stat 4, 4 models, 16
nq=4 , 3 contexts, 16 paired probes
Probe-based certificates
Stat 4, 4 models, 16
n=256 , 3 contexts, 64 probes, nq=8
Forward second differences
Stat 4, 4 models, 16
n=256 , 3 contexts, first query, 4 tangent directions per context, 3 step sizes
Exact Hessian
Stat 4, 5 models, 20
n=256 , 3 contexts, nq=4
Hessian validation
2 datasets, 5 models, 10
n=256
Appendix
Table 5: Grid and replication of every experiment in the paper. Each row names an experiment and gives its grid (the datasets and models it covers) and its replication. A setting is one model evaluated on one dataset. Stat 4 , Core 8 , and Core 7 are the dataset grids defined in the text. n is the number of context rows, nq the number of query points, M the number of label draws used to fit a surrogate, and probes the number of random tangent directions used to estimate a Hessian statistic without forming the full matrix (Appendix B.3 ).
Model
Dataset
ε⋆
cos
rel- L2 med (max)
matched
∥H∥F/∥a∥
TabICL-v2
synthetic_mi
3×10−3
1.0000
0.0008(0.0012)
1.00
5.51
TabICL-v2
kin8nm
3×10−3
1.0000
0.0010(0.0020)
1.00
4.62
TabPFN-v3
synthetic_mi
3×10−3
1.0000
0.0014(0.0046)
1.00
7.29
TabPFN-v3
kin8nm
3×10−3
1.0000
0.0020(0.0063)
1.00
4.99
TabPFN-v2.6
synthetic_mi
3×10−3
1.0000
0.0004(0.0006)
1.00
3.02
TabPFN-v2.6
kin8nm
3×10−3
1.0000
0.0004(0.0005)
1.00
3.14
Appendix
Table 6: Validating the eager-attention Hessians against the native fused-kernel gradient , at n=256 on synthetic_mi and kin8nm , two of the four datasets in the main grid. For each model–dataset setting, ε⋆ is the selected finite-difference step, cos and rel- L2 are the cosine similarity and relative L2 error between the eager Hessian-vector product and the central finite difference of the native gradient along the same probe direction, reported as the median over probes and queries with the maximum in parentheses, matched is the fraction of Hessian columns on which the two computations agree, and ∥H∥F/∥a∥ gives the scale of the Hessian relative to the sensitivity vector for context, not a validation criterion. Every setting meets the thresholds stated in the text.
Model
Dataset
AT/∥PTa∥
AˉT
Att. drop [interval]
Δperm⊥
LS
TabICL-v2
synthetic_mi
5.08
0.991
0.874 [ 0.82,0.93 ]
1.08
0.893
TabICL-v2
kin8nm
4.40
0.993
0.814 [ 0.74,0.86 ]
1.02
0.865
TabICL-v2
pol
35.75
0.972
0.874 [ 0.79,0.92 ]
1.12
0.783
TabICL-v2
superconduct
10.05
0.987
0.749 [ 0.68,0.82 ]
0.99
0.789
TabPFN-v3
synthetic_mi
8.76
0.996
0.933 [ 0.88,0.97 ]
1.01
0.915
TabPFN-v3
kin8nm
4.82
0.993
0.850 [ 0.81,0.91 ]
1.00
0.907
Appendix
Table 7: First-order and probe-based certificate results, by setting. Each row averages three sampled contexts and eight queries per setting at n=256 . The columns AT/∥PTa∥ , AˉT , and Att. drop are estimated from tangent probes (Appendix B.3 ). AT/∥PTa∥ scales the certificate by the norm of the tangent gradient, so that models with larger raw sensitivities do not appear to have a proportionally larger certificate. Att. drop is the relative fall of AT when derivatives cannot pass through the attention scores (Section 4.3 ), and the brackets enclose the three per-context paired-probe 95% bootstrap intervals. Δperm⊥ and LS are the corrected first-order statistics of Appendix A , which likewise vanish for every map that is affine on M .
Att. drop
Model
Dataset
Xˉsep
Range
AˉT
Xsep/∥PTa∥
AT
Xsep
TabICL-v2
synthetic_mi
0.864
[0.836,0.900]
0.990
4.80
0.878
0.895
TabICL-v2
kin8nm
0.772
[0.700,0.813]
0.993
4.17
0.820
0.848
TabICL-v2
pol
0.322
[0.306,0.353]
0.977
20.7
0.885
0.838
TabICL-v2
superconduct
0.566
[0.552,0.575]
0.990
5.94
0.681
0.700
TabPFN-v3
synthetic_mi
0.888
[0.844,0.922]
0.996
8.49
0.935
0.968
Appendix
Table 8: Exact certificate of joint processing, by setting. Behind the per-model means of Table 2 , each row reports Xˉsep for one setting as the mean over three sampled contexts and four queries at n=256 , with the range of the three context means (Range). This certificate is exactly zero for every map that is affine plus one-label-at-a-time on M (Section 3 ), and its float32 implementation floors it at 7×10−11 . AˉT also comes from the exact Hessian, whereas Table 7 estimates it from probes. Xsep/∥PTa∥ scales Xsep by the norm of the tangent gradient. The two Att. drop columns give 1−AT(Hdet)/AT(H) and 1−Xsep(Hdet)/Xsep(H) , the relative falls of AT and Xsep when derivatives cannot pass through the attention scores.
Probe estimate
Exact Hessian
Model
∥PTHPT∥F
AT
AT
Xsep
TabICL-v2
82.1%
82.8%
81.6%
82.0%
TabPFN-v3
87.7%
88.6%
85.9%
90.0%
TabPFN-v2.6
91.5%
91.4%
90.9%
95.3%
TabFM
–
–
78.2%
82.1%
TabDPT
69.6%
70.6%
68.0%
75.3%
Appendix
Table 9: Attention-path norm reductions. Every entry is the relative fall of a norm under the intervention described in the text, not a share of squared curvature and not a loss of accuracy. Values average over four datasets, with three contexts per dataset and four queries per context (eight for the probe estimate of AT ) at n=256 context rows. The probe columns estimate the tangent curvature norm ∥PTHPT∥F from 16 shared probes and AT from 64 probes, and the exact columns use Hessians assembled column by column on all 20 settings. We measured TabFM with the exact Hessian only, so the probe means cover the other four models.
Model
Dataset
Rtask2
Rperm2
RNW-lin2
TabICL-v2
synthetic_mi
0.97
0.68
1.00
TabICL-v2
kin8nm
0.71
0.65
1.00
TabICL-v2
pol
0.68
0.56
1.00
TabPFN-v3
synthetic_mi
0.93
0.28
1.00
TabPFN-v3
kin8nm
0.75
0.55
1.00
TabPFN-v3
pol
0.57
0.13
1.00
Appendix
Table 10: Held-out surrogate fidelity. Held-out R2 , averaged over the six queries, of a surrogate that is affine in y , fitted to each model’s predictions under three label-draw distributions: Rtask2 under the task-preserving resampling described in the text, Rperm2 under label permutations, and RNW-lin2 when the same fitting protocol is applied to a label-linear Nadaraya–Watson control instead of a model. A value of 1 means that the surrogate reproduces the predictions exactly on held-out draws, so the control column shows that the protocol recovers a genuinely affine map whenever one is present. The bottom row gives the median over the 12 settings.
Figure 3: Approximation fidelity and predictive cost. The same 21 settings from three prior-fitted models show (a) the held-out permutation-surrogate R2 , (b) the increase in NMSE when the model’s own predictions are replaced by this surrogate, and (c) the relation between the surrogate’s error and the model’s predictive advantage, which is the held-out NMSE of the best of three tuned label-linear smoothers ( k -nearest neighbors, kernel ridge regression and local-linear regression) minus the model’s own NMSE, so that a negative value means that a smoother predicts better than the model. Replacement increases the error on all 21 settings, but the correlation between the two magnitudes is −0.03 , so this is an approximation comparison and not a causal ablation. TabDPT is not plotted here. A separate permutation fit with 200 label draws at n=96 gives it an R2 of −0.92 and −3.18 on synthetic_mi and kin8nm , against 0.05 and 0.03 under the protocol of Table 10 , so its permutation fidelity depends strongly on the protocol.
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems. They are constructed as transformers that approximate the Bayesian posterior predictive distribution based on a pre-training prior. These univariate predictors can be converted into multivariate ones autoregressively by sampling one target and adding it to the features. However, the faithfulness of the resulting joint has not been investigated. Furthermore, TFMs cannot be evaluated against the posterior itself, at least not on real-world datasets, because the ground-truth distribution is unknown. We therefore propose asking a different question: could a model's predictions result from any joint distribution? To answer this question, we pose two requirements that any such model must satisfy. The first is marginalization consistency, which demands that marginalized conditionals are equal to directly predicted marginals. The second is factorization consistency, which demands that different factorization orders result in equal joint distributions. Every TFM that we evaluate violates both of these requirements for both classification and regression across all datasets.
Christian Klötergens, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme +1
Institute of Computer Science & VWFS DARC, University of Hildesheim, Hildesheim, Germany
Tabular foundation models with different architectures converge in accuracy across a range of classification and regression tasks. This raises questions a leaderboard cannot answer: (i) whether the models execute the same in-context algorithm, (ii) where row, column, and class-permutation invariances originate, and (iii) how robust they are under perturbations engineered against the inferred mechanism. We characterize all three. The model families realize qualitatively distinct similarity-based readouts: from an attention-weighted vote over context labels to a class-conditional mean readout, each confirmed by causal intervention. We find that the representation collapse highlighted in prior work is not a practical concern for them. Each model's permutation invariances trace to specific positional parameters whose removal preserves accuracy and makes approximate invariance exact. Perturbations engineered against each readout reproduce predicted failure modes; hub and rank attacks isolate them from refit baselines. Together these results give a mechanistic account of contemporary tabular foundation models and identify which inductive biases govern both their accuracy and characteristic failures.
Marin Biloš, James T. Wilson, Anderson Schneider +1