Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.
Figures & tables
Figure 1: Retro lies on the Pareto frontier of three mainstream benchmarks.
Figure 2: A housing example of complementary neighborhood views and retrospective integration. Top: a person considers size, location, and condition, then revisits these assessments to reach a decision. Bottom: the TFM schematic places the same homes in different neighborhoods across representations and integrates earlier information in the final representation. This illustrates a possible organization of tabular inference.
Figure 3: Layer-wise changes in predictive margin for TabICLv2 and a retrospective variant on SDSS17. The left and right grids show query representations across all 12 layers of TabICLv2 and our retrospective variant, respectively. Within each model, representations are projected onto a fixed PCA basis fitted to its final-layer support embeddings; coordinates are therefore comparable across layers within a model, but not between models. Color indicates the change in true-class probability margin from the preceding layer ( Δm ), with green denoting an increase and purple a decrease. Circles mark queries changing from incorrect to correct, and crosses mark the reverse. L1 is shown in gray because no preceding layer is measured. Insets enlarge L1, L4, and L7; the L4 and L7 insets use separately labeled, narrower color scales to reveal small early-layer changes. TabICLv2 exhibits limited directly readable changes through much of its depth, whereas the retrospective variant shows changes for different queries across more stages. Full four-model visualizations and aggregate results over 38 datasets appear in Sections C.1 and A.4 .
Figure 4: Sample-state flows across depth. Frozen-readout trajectories for TabICLv2 (left) and our ungated retrospective variant (right), averaged equally over 38 classification datasets. Correct merges always-correct and previously-wrong but currently-correct queries; the two wrong states distinguish whether a query was ever correct. Ribbon widths show query fractions across all 12 layers. The displayed average measures correctness switches over L1 → L2 through L10 → L11; only this summary excludes the final transition.
Figure 5
Figure 7: The Retro architecture. A table encoder produces row representations for a retrospective ICL stack. Within each layer, separate aggregation modules construct attention and feed-forward inputs from available historical contributions. An input-conditioned gate modulates attention outputs before projection. Completed two-layer update groups remain available to subsequent layers and the final aggregation used for prediction.
Figure 8: Elo rankings across three benchmarks. Each panel shows the five highest-rated methods. Candidate sets and scoring protocols differ, so ratings should be compared within panels. Complete rankings appear in Appendix E .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Any revision (%)
Revision / step (%)
E∣Δm∣
Early motion (%)
TabICLv2
63.8
9.82
0.197
15.6
TabPFN-3
66.8
7.55
0.149
48.5
EXAONE-Tabular
71.8
16.12
0.319
38.9
Ours
79.5
24.66
0.488
60.7
Appendix
Table 1: Descriptive frozen-readout dynamics, equally averaged over 38 datasets (one official split each). Bold marks the largest value in each column. Early motion is the share of absolute probability-margin change before the final third of depth. TabPFN-3 has 24 layers; the other models have 12. Native-step averages complement the depth-dependent any-revision statistic, but do not equate compute or depth granularity across architectures.
Figure 9: TabICLv2 on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 10: Our retrospective model on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 11: TabPFN-3 on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 12: EXAONE-Tabular on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 13: Four-model sample-state flows on SDSS17. Node heights and ribbon widths are percentages of all outer queries. Correct merges both currently-correct histories; the two wrong states distinguish whether a query was ever correct at an earlier layer. All native layers are retained.
Figure 14: Four-model sample-state flows averaged over 38 classification datasets. Node heights and ribbon widths are equal-weight means of per-dataset query percentages, using the same three states as Figure 4 . All native layers are shown. The average inter-layer prediction change excludes only each model’s final transition from the summary, not from the plotted flows.
Figure 15: TabICLv2 on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 16: Our retrospective model on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 17: TabPFN-3 on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 18: EXAONE-Tabular on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 19: TabICLv2 on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 20: Our retrospective model on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 21: TabPFN-3 on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 22: EXAONE-Tabular on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 23: Sensitivity to update direction and magnitude. Top: mean degradation relative to learned gating, with 95% dataset-bootstrap intervals. Bottom: dataset-level effects; points above the diagonal indicate greater harm from replacing direction. The classification inset enlarges the region near zero. Positive values mean lower accuracy or higher NRMSE. Replacement distances are not matched.
Figure 24: Neighborhood-constrained versus random gate exchange. Each point is a dataset, averaging three exchange seeds. Below the diagonal, exchange within support-derived neighborhoods is less harmful. Group sizes and moved-row counts are matched, within-neighborhood exchanges also induce smaller gate displacement.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
1882.9
109.9
99.5
4.31
EXAONE-Tabular (default)
1849.7
72.8
49.7
4.89
Retro
1741.6
84.3
72.5
7.26
TabPFN-3 (default)
1728.9
76.1
55.1
7.58
Xiaomi-TabLDM (default)
1673.7
74.2
66.0
9.11
TabPFN-2.6 (default)
1669.5
62.7
44.9
9.24
Appendix
Table 2: TabArena: overall ranking. All 37 selected methods on 51 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
1833.3
118.3
107.4
4.76
EXAONE-Tabular (default)
1821.1
82.0
56.0
5.00
TabPFN-3 (default)
1690.7
74.8
68.0
8.12
Retro
1680.4
79.1
70.8
8.42
TabPFN-2.6 (default)
1638.4
56.6
50.2
9.70
Xiaomi-TabLDM (default)
1631.5
70.7
60.1
9.93
Appendix
Table 3: TabArena: classification ranking. All 38 selected methods on 38 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
2165.6
254.1
73.4
3.56
Retro
2082.1
332.6
117.5
4.83
EXAONE-Tabular (default)
2043.9
251.0
83.9
5.52
TabPFN-3 (default)
1966.4
320.5
122.8
7.15
Xiaomi-TabLDM (default)
1928.5
348.6
176.1
8.05
Nori-30M (default)
1895.0
254.2
76.1
8.90
Appendix
Table 4: TabArena: regression ranking. All 39 selected methods on 13 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1856.0
338.5
339.8
3.99
Retro
1755.4
280.9
443.1
5.41
EXAONE-Tabular
1745.2
254.4
239.3
5.22
Xiaomi-TabLDM
1717.2
223.6
290.2
6.05
TabPFN-3
1697.7
357.0
347.5
5.68
TabICLv2
1679.3
218.9
384.4
6.22
Appendix
Table 5: TALENT: overall ranking. All 17 selected methods on 300 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1829.8
306.9
408.4
4.11
Retro
1719.4
312.6
299.9
5.89
TabICLv2
1715.2
316.5
362.1
5.93
EXAONE-Tabular
1708.3
238.5
255.4
5.53
TabPFN-3
1676.8
267.1
228.3
5.80
Xiaomi-TabLDM
1670.8
304.8
295.4
6.37
Appendix
Table 6: TALENT: classification ranking. All 17 selected methods on 200 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1886.2
254.1
281.8
3.76
Retro
1873.7
172.4
231.4
4.44
EXAONE-Tabular
1868.3
290.2
357.7
4.60
Xiaomi-TabLDM
1780.7
210.6
257.1
5.42
TabPFN-3
1729.9
315.9
461.8
5.45
TabICLv2
1666.1
293.8
274.2
6.80
Appendix
Table 7: TALENT: regression ranking. All 17 selected methods on 100 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
Retro
1797.0
78.4
60.4
2.90
KurveRSC
1781.2
92.8
93.6
3.05
TabPFN-rel-local
1721.8
116.2
80.2
3.62
GraphSAGE
1666.4
90.8
78.6
4.19
RelGT
1575.2
126.8
103.3
5.17
RDBLearn
1555.1
89.6
93.1
5.38
Appendix
Table 8: RelArena: overall ranking. All 10 selected methods on 21 tasks.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
Retro
2185.9
266.1
45.3
3.08
KurveRSC
2151.6
340.0
111.5
3.42
TabPFN-rel-local
2135.2
315.9
122.3
3.58
GraphSAGE
2103.2
358.6
135.7
3.92
RelGT
1981.0
352.6
124.4
5.25
RelGNN-ES
1958.1
292.1
106.1
5.50
Appendix
Table 9: RelArena: classification ranking. All 10 selected methods on 12 tasks.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
KurveRSC
1736.4
321.6
143.4
2.56
Retro
1722.3
302.3
130.4
2.67
TabPFN-rel-local
1609.0
247.2
169.5
3.67
GraphSAGE
1519.0
103.4
94.2
4.56
RelGT
1469.8
208.6
185.5
5.06
RDBLearn
1453.4
158.5
188.5
5.22
Appendix
Table 10: RelArena: regression ranking. All 10 selected methods on 9 tasks.
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
What reusable computation should a tabular foundation model learn when every table defines a new supervised task? We develop in-situ representation refinement: support labels guide updates to the episode's representations, and these updates transfer to unlabeled queries without changing model parameters. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-free contextual stack with selected low-rank feature interaction and typed memory. RefineICL-L24 reaches 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29. A benchmark-informed continuation reaches 1644.8 Elo on the 38-dataset TabArena snapshot, 31.4 Elo above TabPFN-3 under the same evaluation. It also improves all four reported metrics over TabPFN-v3 on both TabZilla views. In a matched 100K-update depth grid, an expanded FFN gives no consistent validation benefit and uses 60.2% more peak inference memory at L8. Internal interventions show that support representations are more than a static source of labels: removing one intermediate support update, while preserving the query output, increases final query cross-entropy in all 72 tested episodes. Together, the derivation and interventions explain how attention-gated updates can construct a task-specific predictor in context.
Tabular foundation models (TFMs), such as TabPFN-2.6, TabICLv2, ConTextTab, Mitra, LimiX, and TabDPT, achieve strong zero-shot performance through in-context learning, but their inductive biases remain fixed at inference time. Adapting a pretrained TFM to a specific dataset or task typically requires either full fine-tuning, which is computationally expensive, or parameter-efficient tuning methods (PEFT) such as LoRA, which must be tailored to the internal architecture of each TFM. Furthermore, the evidence on whether weight-space fine-tuning improves accuracy or calibration is mixed \citep{tanna_exploring_2026,rubachev_finetuning_2025}. We introduce TFM-Retouche, a lightweight input-space residual adapter that is architecture-agnostic by design with respect to the frozen TFM backbone. TFM-Retouche learns a small residual correction in the input space to align the input data with the inductive biases of the pretrained model. The adapter is trained end-to-end through the frozen TFM, with a post-training identity guard that falls back to the unmodified TFM whenever adaptation does not help on held-out validation. On TabArena-Lite (51 datasets spanning binary classification, multiclass classification, and regression), TabICLv2-Retouche -- the framework instantiated on TabICLv2 -- is the top-ranked method on the leaderboard with light per-task tuning and ensembling, lifting aggregate Elo by +56 over the frozen TabICLv2 base and sitting on the Pareto front of predictive quality versus both training and inference time.