Tabular prediction can benefit from in-table rows as few-shot evidence, yet existing tabular models typically perform instance-wise inference and LLM-based prompting is often brittle. Models do not consistently leverage relevant rows, and noisy context can degrade performance. To address this challenge, we propose TabSieve, a select-then-predict framework that makes evidence usage explicit and auditable. Given a table and a query row, TabSieve first selects a small set of informative rows as evidence and then predicts the missing target conditioned on the selected evidence. To enable this capability, we construct TabSieve-SFT-40K by synthesizing high-quality reasoning trajectories from 331 real tables using a strong teacher model with strict filtering. Furthermore, we introduce TAB-GRPO, a reinforcement learning recipe that jointly optimizes evidence selection and prediction correctness with separate rewards, and stabilizes mixed regression and classification training via dynamic task-advantage balancing. Experiments on a held-out benchmark of 75 classification and 52 regression tables show that TabSieve consistently improves performance across shot budgets, with average gains of 2.92% on classification and 4.45% on regression over the second-best baseline. Further analysis indicates that TabSieve concentrates more attention on the selected evidence, which improves robustness to noisy context.
Figures & tables
Figure 1: We contrast three prediction paradigms. Traditional models cannot explicitly interpret how context is used. Strong LLMs with in-context prompting can be distracted by noisy context . TabSieve first performs evidence selection and then conducts noise-filtered reasoning to produce the final prediction .
Figure 2: Explicit evidence selection concentrates attention on informative rows.
Figure 3: Effect of explicit row selection under evidential or noisy in-table context.
Figure 4: Training pipeline of TabSieve. We synthesize select-then-predict trajectories from a teacher model to build SFT dataset. Starting from the SFT-initialized model, TAB-GRPO extends GRPO with task-advantage balancing to mitigate optimization imbalance and strengthen context selection for robust in-context learning.
Classification (Accuracy ↑ )
Regression (NMAE ↓ )
Model
0
4
8
16
32
Rank z
Rank i
0
4
8
16
32
Rank z
Rank i
Traditional Tabular Models
XGBoost Chen and Guestrin (2016)
–
49.93
51.79
58.38
62.98
–
10.63
–
0.237
0.200
0.184
0.160
–
8.23
CatBoost Prokhorenkova et al. (2019)
–
51.25
56.42
65.88
72.64
–
8.95
–
0.235
0.192
0.179
0.159
–
7.56
FT-Transformer Gorishniy et al. (2023)
–
54.31
59.74
68.60
71.45
–
8.37
–
0.363
0.363
0.357
0.354
–
11.56
TabPFN Hollmann et al. (2023)
–
52.68
60.70
69.15
73.67
–
8.09
–
0.228
0.186
0.165
0.144
–
6.70
Table 1: Classification (Accuracy ↑ ) and Regression (NMAE ↓ ) results averaged over all datasets under zero-shot and few-shot settings. Rank z denotes the average rank in the zero-shot setting; Rank i is averaged across all datasets and all few-shot settings. Δ denotes the performance gain relative to Qwen3-8B.
Figure 5: Ablation on training stages and RL components.
Figure 6: Training reward curves of TAB-GRPO and Vanilla GRPO.
Figure 7: Ablation on evidence selection.
Figure 8: Case study of evidence selection.
Figure 9: Data distribution of the SFT dataset. Left: dataset composition across the five shot budgets, with each slice further decomposed into classification and regression trajectories. Right: distribution of the evidence ratio for K>0 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameters
Value
Epochs
4.0
Micro batch size
4
Gradient accumulation
4
Learning rate
1×10−4
Warmup ratio
0.1
LR scheduler
Cosine
Appendix
Table 2: Core SFT hyperparameters for cold-start.
Hyperparameters
Value
Epochs
2.0
Train batch size
256
Rollout sampling num
8
Rollout temperature
1.0
KL coefficient
0.001
Learning rate
1×10−6
Appendix
Table 3: Core RL hyperparameters for TAB-GRPO.
Figure 10: Training and inference prompt templates of the TabSieve.
Figure 11: Inference prompt templates of the general LLMs.
Evidence
4
8
16
32
∣Einf∣
4592
8594
17791
34313
∣Eemb∣
5681
10327
23717
46235
∣Eemb∩Eiinf∣
3551
6816
14968
28567
∣Einf∣∣Eemb∩Einf∣
77.33%
79.31%
84.13%
83.25%
Appendix
Table 4: Overlap between embedding-based evidence and influence-based evidence under different shot budgets.
Evidence
0
4
8
16
32
Qwen3-8B
48.18
58.77
60.91
63.01
65.67
Influence-based
54.41
65.23
69.06
70.72
74.42
Embedding-based
54.63
65.77
69.32
70.98
74.33
Appendix
Table 5: Classification results of different evidence construction strategies. We report accuracy (↑) averaged over all classification datasets under all shot settings.
Evidence
0
4
8
16
32
Qwen3-8B
0.325
0.268
0.239
0.235
0.222
Influence-based
0.262
0.184
0.171
0.149
0.132
Embedding-based
0.266
0.192
0.169
0.158
0.139
Appendix
Table 6: Regression results of different evidence construction strategies. We report NMAE (↓) averaged over all regression datasets under all shot settings.
Model
Acc. ↑
NMAE ↓
Qwen2.5-72B w. Std. Prompt
65.11
0.217
Qwen2.5-72B w. TabSieve Prompt
65.92
0.209
Qwen3-8B w. Std. Prompt
59.31
0.258
Qwen3-8B w. TabSieve Prompt
59.79
0.251
TabSieve (Ours)
67.01
0.185
Appendix
Table 7: Prompt ablation. We report average classification accuracy and regression NMAE over all shot settings.
Hyperparameter
Value
Acc. ↑
NMAE ↓
G
4
65.97
0.209
8
67.01
0.185
12
66.82
0.179
ηmin
0.7
66.57
0.197
0.8
67.01
0.185
0.9
66.14
0.192
Appendix
Table 8: Sensitivity analysis of key RL hyperparameters.
Variant
Acc. ↑
NMAE ↓
w/o EMA Smoothing
65.12
0.207
w/o ηmin
66.04
0.202
TabSieve
67.01
0.185
Appendix
Table 9: Component-level ablation of the task-advantage balancing mechanism in TAB-GRPO. We report average classification accuracy and regression NMAE over all shot settings.
Model
0
4
8
16
32
DeepSeek-R1-70B
41.06
52.21
54.75
53.12
59.93
GPT-OSS-20B
40.75
48.99
50.15
52.48
53.87
Qwen3-8B
26.76
34.62
43.22
47.74
50.92
TabSieve
43.19
54.10
56.06
55.92
60.31
Appendix
Table 10: Macro-F1 ( ↑ ) averaged over all classification datasets under zero-shot and few-shot settings. The best result in each column is in bold and the second best is underlined.
Variant
Acc. ↑
NMAE ↓
Qwen3-8B
59.31
0.258
Qwen3-8B w/o Selected Rows
43.57
0.336
Appendix
Table 11: Utility of the rows selected by TabSieve. We remove the selected rows from the context and re-evaluate Qwen3-8B on the remaining rows. Results are averaged over all shot settings.
Figure 12: Zero-shot case in classification tasks. TabSieve first analyzes the semantic meaning of the patient attributes and then predicts the label based on medical knowledge about stroke risk.
Figure 13: Zero-shot case in regression tasks. TabSieve recognizes that the target is the invariant mass of a two-electron collision system, and then derives the final prediction by applying the physics formula.
Figure 14: Few-shot case. TabSieve follows a select-then-predict procedure. It first analyzes the table semantics and the influence of features on the target, then selects the evidence rows, and finally performs prediction mainly based on these selected rows.
We introduce TabPFN-3.5, our new flagship Tabular Foundation Model. It significantly outperforms its predecessor, TabPFN-3, and all existing baselines across a broad range of tabular problems. TabPFN-3.5 sets a new state of the art on standard tabular prediction in TabArena, and extends it to the data practitioners encounter in practice: non-i.i.d. data with temporal or grouped splits, tables with strings, text and images, high-cardinality categorical features, and wide tables with many features. These gains carry over to our task-specific harnesses: state of the art on relational data and stronger time-series forecasting. For faster inference, our variant TabPFN-3.5-Fast runs up to 3x faster than TabPFN-3 while keeping most of the accuracy gains. In addition, we upgrade TabPFN-3.5-Plus, expanding our multimodal capabilities with advanced text and date handling alongside proprietary inference optimizations. Finally, we release a new version of our Thinking mode, TabPFN-3.5-Thinking, which scales inference-time computation to push the state of the art further. It benefits from our stronger base model and from inference-time improvements that make it up to 12x faster than TabPFN-3-Thinking.
We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders. Unlike existing tabular in-context learners, which entangle feature representations with a specific prediction target, our design produces target-agnostic row embeddings that can be leveraged across a wide range of downstream tasks within a table-native in-context learning setup. We demonstrate this flexibility on six distinct problems: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. Both the encoder and the task-specific decoders are trained on a large corpus of real-world, unlabeled tables. FlexTab achieves state-of-the-art performance on classification, regression, anomaly detection and entity matching, while remaining competitive with specialized models on entity classification in a relational setting. These results demonstrate that a single shared encoder, paired with task-specific decoders, can serve as an effective general-purpose backbone for diverse tabular prediction problems. The inference code and checkpoints will be made publicly available at https://github.com/SAP-samples/flextab.
Marek Polewczyk, Maximilian Schambach, Marco Spinaci +2
Tabular foundation models, exemplified by TabPFN, perform prediction via in-context learning, inferring test labels directly from labeled training examples. They have demonstrated competitive performance, particularly on small-to-medium datasets. However, recent tabular foundation models often improve accuracy with increasingly complex architectures, incurring higher inference cost and limiting practical deployment. In this work, we revisit the original TabPFN design and show that a lightweight row-wise attention-only backbone can remain highly competitive with two simple enhancements: a gated attention stabilization mechanism and a small set of learnable register tokens that provide global context and improve pretraining quality. The resulting model, TabSwift, supports both classification and regression, and is competitive with stronger tabular foundation models (e.g., TabPFN v2 and TabICL) while being more efficient at inference. For latency-sensitive serving, we further introduce an adaptive layer-wise early-exit mechanism that dynamically adjusts inference depth per sample. Overall, TabSwift enables efficient and anytime tabular in-context learning for practical deployments.
Si-Yang Liu, Han-Jia Ye
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China.