Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.
Figures & tables
Figure 1: Retro lies on the Pareto frontier of three mainstream benchmarks.
Figure 2: A housing example of complementary neighborhood views and retrospective integration. Top: a person considers size, location, and condition, then revisits these assessments to reach a decision. Bottom: the TFM schematic places the same homes in different neighborhoods across representations and integrates earlier information in the final representation. This illustrates a possible organization of tabular inference.
Figure 3: Layer-wise changes in predictive margin for TabICLv2 and a retrospective variant on SDSS17. The left and right grids show query representations across all 12 layers of TabICLv2 and our retrospective variant, respectively. Within each model, representations are projected onto a fixed PCA basis fitted to its final-layer support embeddings; coordinates are therefore comparable across layers within a model, but not between models. Color indicates the change in true-class probability margin from the preceding layer ( Δm ), with green denoting an increase and purple a decrease. Circles mark queries changing from incorrect to correct, and crosses mark the reverse. L1 is shown in gray because no preceding layer is measured. Insets enlarge L1, L4, and L7; the L4 and L7 insets use separately labeled, narrower color scales to reveal small early-layer changes. TabICLv2 exhibits limited directly readable changes through much of its depth, whereas the retrospective variant shows changes for different queries across more stages. Full four-model visualizations and aggregate results over 38 datasets appear in Sections C.1 and A.4 .
Figure 4: Sample-state flows across depth. Frozen-readout trajectories for TabICLv2 (left) and our ungated retrospective variant (right), averaged equally over 38 classification datasets. Correct merges always-correct and previously-wrong but currently-correct queries; the two wrong states distinguish whether a query was ever correct. Ribbon widths show query fractions across all 12 layers. The displayed average measures correctness switches over L1 → L2 through L10 → L11; only this summary excludes the final transition.
Figure 5
Figure 7: The Retro architecture. A table encoder produces row representations for a retrospective ICL stack. Within each layer, separate aggregation modules construct attention and feed-forward inputs from available historical contributions. An input-conditioned gate modulates attention outputs before projection. Completed two-layer update groups remain available to subsequent layers and the final aggregation used for prediction.
Figure 8: Elo rankings across three benchmarks. Each panel shows the five highest-rated methods. Candidate sets and scoring protocols differ, so ratings should be compared within panels. Complete rankings appear in Appendix E .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Any revision (%)
Revision / step (%)
E∣Δm∣
Early motion (%)
TabICLv2
63.8
9.82
0.197
15.6
TabPFN-3
66.8
7.55
0.149
48.5
EXAONE-Tabular
71.8
16.12
0.319
38.9
Ours
79.5
24.66
0.488
60.7
Appendix
Table 1: Descriptive frozen-readout dynamics, equally averaged over 38 datasets (one official split each). Bold marks the largest value in each column. Early motion is the share of absolute probability-margin change before the final third of depth. TabPFN-3 has 24 layers; the other models have 12. Native-step averages complement the depth-dependent any-revision statistic, but do not equate compute or depth granularity across architectures.
Figure 9: TabICLv2 on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 10: Our retrospective model on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 11: TabPFN-3 on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 12: EXAONE-Tabular on SDSS17. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 13: Four-model sample-state flows on SDSS17. Node heights and ribbon widths are percentages of all outer queries. Correct merges both currently-correct histories; the two wrong states distinguish whether a query was ever correct at an earlier layer. All native layers are retained.
Figure 14: Four-model sample-state flows averaged over 38 classification datasets. Node heights and ribbon widths are equal-weight means of per-dataset query percentages, using the same three states as Figure 4 . All native layers are shown. The average inter-layer prediction change excludes only each model’s final transition from the summary, not from the plotted flows.
Figure 15: TabICLv2 on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 16: Our retrospective model on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 17: TabPFN-3 on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 18: EXAONE-Tabular on superconductivity. All native layers in the fixed final-support PCA basis. Green/purple indicate reductions/increases in absolute error under the frozen final-support Ridge readout, normalized by the support-target standard deviation. All four models share the same color scale. The first layer is gray; no binary correctness states are assigned.
Figure 19: TabICLv2 on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 20: Our retrospective model on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 21: TabPFN-3 on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 22: EXAONE-Tabular on Is-this-a-good-customer. All native layers in the fixed final-support PCA basis. Colors show adjacent probability-margin change under the frozen readout; rings and crosses mark corrections and deteriorations. The first layer is gray. Extraction and display conventions are given in Sections A.1 and A.3 .
Figure 23: Sensitivity to update direction and magnitude. Top: mean degradation relative to learned gating, with 95% dataset-bootstrap intervals. Bottom: dataset-level effects; points above the diagonal indicate greater harm from replacing direction. The classification inset enlarges the region near zero. Positive values mean lower accuracy or higher NRMSE. Replacement distances are not matched.
Figure 24: Neighborhood-constrained versus random gate exchange. Each point is a dataset, averaging three exchange seeds. Below the diagonal, exchange within support-derived neighborhoods is less harmful. Group sizes and moved-row counts are matched, within-neighborhood exchanges also induce smaller gate displacement.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
1882.9
109.9
99.5
4.31
EXAONE-Tabular (default)
1849.7
72.8
49.7
4.89
Retro
1741.6
84.3
72.5
7.26
TabPFN-3 (default)
1728.9
76.1
55.1
7.58
Xiaomi-TabLDM (default)
1673.7
74.2
66.0
9.11
TabPFN-2.6 (default)
1669.5
62.7
44.9
9.24
Appendix
Table 2: TabArena: overall ranking. All 37 selected methods on 51 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
1833.3
118.3
107.4
4.76
EXAONE-Tabular (default)
1821.1
82.0
56.0
5.00
TabPFN-3 (default)
1690.7
74.8
68.0
8.12
Retro
1680.4
79.1
70.8
8.42
TabPFN-2.6 (default)
1638.4
56.6
50.2
9.70
Xiaomi-TabLDM (default)
1631.5
70.7
60.1
9.93
Appendix
Table 3: TabArena: classification ranking. All 38 selected methods on 38 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM (default)
2165.6
254.1
73.4
3.56
Retro
2082.1
332.6
117.5
4.83
EXAONE-Tabular (default)
2043.9
251.0
83.9
5.52
TabPFN-3 (default)
1966.4
320.5
122.8
7.15
Xiaomi-TabLDM (default)
1928.5
348.6
176.1
8.05
Nori-30M (default)
1895.0
254.2
76.1
8.90
Appendix
Table 4: TabArena: regression ranking. All 39 selected methods on 13 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1856.0
338.5
339.8
3.99
Retro
1755.4
280.9
443.1
5.41
EXAONE-Tabular
1745.2
254.4
239.3
5.22
Xiaomi-TabLDM
1717.2
223.6
290.2
6.05
TabPFN-3
1697.7
357.0
347.5
5.68
TabICLv2
1679.3
218.9
384.4
6.22
Appendix
Table 5: TALENT: overall ranking. All 17 selected methods on 300 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1829.8
306.9
408.4
4.11
Retro
1719.4
312.6
299.9
5.89
TabICLv2
1715.2
316.5
362.1
5.93
EXAONE-Tabular
1708.3
238.5
255.4
5.53
TabPFN-3
1676.8
267.1
228.3
5.80
Xiaomi-TabLDM
1670.8
304.8
295.4
6.37
Appendix
Table 6: TALENT: classification ranking. All 17 selected methods on 200 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
TabFM
1886.2
254.1
281.8
3.76
Retro
1873.7
172.4
231.4
4.44
EXAONE-Tabular
1868.3
290.2
357.7
4.60
Xiaomi-TabLDM
1780.7
210.6
257.1
5.42
TabPFN-3
1729.9
315.9
461.8
5.45
TabICLv2
1666.1
293.8
274.2
6.80
Appendix
Table 7: TALENT: regression ranking. All 17 selected methods on 100 datasets.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
Retro
1797.0
78.4
60.4
2.90
KurveRSC
1781.2
92.8
93.6
3.05
TabPFN-rel-local
1721.8
116.2
80.2
3.62
GraphSAGE
1666.4
90.8
78.6
4.19
RelGT
1575.2
126.8
103.3
5.17
RDBLearn
1555.1
89.6
93.1
5.38
Appendix
Table 8: RelArena: overall ranking. All 10 selected methods on 21 tasks.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
Retro
2185.9
266.1
45.3
3.08
KurveRSC
2151.6
340.0
111.5
3.42
TabPFN-rel-local
2135.2
315.9
122.3
3.58
GraphSAGE
2103.2
358.6
135.7
3.92
RelGT
1981.0
352.6
124.4
5.25
RelGNN-ES
1958.1
292.1
106.1
5.50
Appendix
Table 9: RelArena: classification ranking. All 10 selected methods on 12 tasks.
Method
Elo ↑
Elo+
Elo −
Average Rank ↓
KurveRSC
1736.4
321.6
143.4
2.56
Retro
1722.3
302.3
130.4
2.67
TabPFN-rel-local
1609.0
247.2
169.5
3.67
GraphSAGE
1519.0
103.4
94.2
4.56
RelGT
1469.8
208.6
185.5
5.06
RDBLearn
1453.4
158.5
188.5
5.22
Appendix
Table 10: RelArena: regression ranking. All 10 selected methods on 9 tasks.
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany