Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.
Figures & tables
Figure 1: TICDA overview. TICDA (1) extracts TFM embeddings, (2) trains a linear surrogate on the classification task of interest, (3) estimates influence score based on the surrogate and (4) can be applied for label-error detection, context curation, cross-TFM transfer and active learning.
LOO
DemoShapley
IG
DETAIL
TICDA (ours)
TabDPT
Faith. ↑
—
0.91∗∗∗±0.07
0.20±0.35
−0.06±0.18
0.79±0.09
AUC ↑
0.90∗∗∗±0.08
0.88±0.10
0.60±0.20
0.47±0.13
0.88±0.08
TabICL
Faith. ↑
—
0.65±0.15
0.30±0.22
−0.11±0.22
0.67±0.21
AUC ↑
0.80±0.16
0.76±0.11
0.64±0.18
0.37±0.17
0.89∗∗∗±0.08
Table 1: Data attribution for in-context demonstrations on tabular foundation models. Results are aggregated over 38 datasets (mean ± std). Best results are bold , second best underlined ; stars compare the winner with the runner-up (two-sided paired t -test, ∗p<.10 , ∗∗p<.05 , ∗∗∗p<.01 ).
Figure 2: Average computation time per method.
Configuration
Metrics
Influence
Dimension
Faith. ↑
AUC ↑
Time (s) ↓
query
20%
0.32±0.16
0.85±0.10
4.90±1.30
query
full
0.33±0.16
0.86±0.10
5.09±1.31
self
20%
0.67±0.21
0.89±0.08
3.51±0.68
self
full
0.67±0.21
0.89±0.08
3.71±0.69
Table 2: Ablation study on TabICL (38 datasets, 20% label corruption, mean ± std).
Figure 3: Curation: aggregated balanced accuracy evolution on TabDPT and TabICL under different levels of demonstration corruption. Each method ranks in-context demonstrations by its attribution scores, and the lowest-scored demonstrations are removed.
Demonstrations removed
Target
Source
0%
10%
20%
50%
TabPFN-3
Random
59.0±14.9
58.7±14.6
58.3±14.4
56.6±14.1
TabICLv2
59.0±14.9
63.5±15.8
65.0±15.8
64.4±15.4
TabDPT
59.0±14.9
62.6±15.7
64.0±15.7
63.4±15.2
TabPFN-3.5
Random
60.4±15.5
60.1±15.2
59.8±15.1
57.9±14.3
TabICLv2
60.4±15.5
64.1±15.9
64.5±15.9
64.6±15.5
Table 3: TICDA TDA Transfer for context curation between source and target models. Balanced accuracy (mean ± std, aggregated over all datasets) is reported at several levels of removal.
Figure 7
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Demonstrations in the context
Remaining demonstrations
Queries
512
2.90 ± 0.60
2.93 ± 0.62
1,024
1.73 ± 0.61
1.71 ± 0.61
2,048
0.89 ± 0.24
0.88 ± 0.25
4,096
0.48 ± 0.14
0.48 ± 0.15
8,192
0.24 ± 0.07
0.24 ± 0.08
Appendix
Table 5: Relative change ρij of the TabICL representations when one demonstration is removed (Equation 17 ), in percent of the norm of the representation, for contexts of increasing size. For each dataset, we take the median of ρij over all pairs of a removed demonstration i and a row j , and we report the mean ± standard deviation of this median over the four datasets. A value of zero would mean that the representations do not change.
Figure 5: Relative change ρij of the TabICL representations when one demonstration is removed, against the number n of demonstrations in the context, on logarithmic axes. The markers show the mean over the four datasets of the per-dataset median, as in Table 5 , and the band one standard deviation for the remaining demonstrations. The change of the queries coincides with that of the remaining demonstrations, and both follow the dotted line, which decreases as 1/n .