Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark
Figures & tables
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Time ( \SIUnitSymbolMicros )
Bandwidth ( GB/s )
R
C
Layout A
Layout B
Layout A
Layout B
Ratio A/B
1024
20
65.0
64.9
1290.6
1293.1
1.002
2048
50
310.3
310.8
1351.6
1349.4
0.998
4096
50
619.5
619.7
1354.1
1353.6
1.000
8192
100
2479.5
2481.8
1353.3
1352.0
0.999
8192
16
396.5
397.6
1354.0
1350.1
0.997
Appendix
Table 1 : Microbenchmark of .contiguous() for the row-first layout (Layout A) (B,R,C,H,D) vs. the column-first layout (Layout B) (B,C,R,H,D) with B=2 , H=8 , D=64 . The transpose(1,2) call makes the non-contiguous dimension contiguous. Times in \SIUnitSymbolMicros ; bandwidth in GB/s . All ratios are within 0.3 % of unity, confirming that layout choice does not affect copy cost.
N
Total size ( MB )
Time ( \SIUnitSymbolMicros )
Bandwidth ( GB/s )
32
9
24.1
783
64
18
33.2
1138
128
36
56.8
1330
256
72
115.4
1308
512
144
220.5
1370
1024
288
422.7
1429
Appendix
Table 2 : Measured .contiguous() bandwidth for three copies (Q, K, V) on the H100 at the row attention shapes used in the benchmark. Configuration: Beff=64,H=12,N,D=64 in bfloat16 .
Tabular foundation models, exemplified by TabPFN, perform prediction via in-context learning, inferring test labels directly from labeled training examples. They have demonstrated competitive performance, particularly on small-to-medium datasets. However, recent tabular foundation models often improve accuracy with increasingly complex architectures, incurring higher inference cost and limiting practical deployment. In this work, we revisit the original TabPFN design and show that a lightweight row-wise attention-only backbone can remain highly competitive with two simple enhancements: a gated attention stabilization mechanism and a small set of learnable register tokens that provide global context and improve pretraining quality. The resulting model, TabSwift, supports both classification and regression, and is competitive with stronger tabular foundation models (e.g., TabPFN v2 and TabICL) while being more efficient at inference. For latency-sensitive serving, we further introduce an adaptive layer-wise early-exit mechanism that dynamically adjusts inference depth per sample. Overall, TabSwift enables efficient and anytime tabular in-context learning for practical deployments.
Si-Yang Liu, Han-Jia Ye
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China.
With the recent rise and adoption of tabular foundation models, optimizing their inference performance becomes an emerging field for efficiency research. While the models are architecturally similar to transformer-based large language models (LLMs), the size and serving patterns differ significantly. We show that the focus should be on the attention calculation and less on weight or KV cache quantization, which are more popular in LLMs. We develop a quantization strategy for queries, keys, and values to FP8 and use explicit FP8 matrix multiplication instructions to speed up the attention calculation. We find that it is crucial to align the quantization error in the test rows with the quantization error in the training rows, as otherwise the accuracy drops drastically. Our Triton kernel achieves a speedup up to 1.7x over regular 16-bit kernels, and we show that on TabPFN-v3 and TabICLv2 there is no relevant accuracy loss across TabArena and BeyondArena.
Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. However, recent approaches with either cell-based architectures or retrieval have sacrificed efficiency for raw performance, restricting their utility in situations where compute is limited or inference speed is crucial. We adopt an alternate approach, sticking with row-based attention while incorporating long context pre-training to eliminate the need for retrieval. By combining this with architectural improvements and SSL pre-training on a newly-sourced, larger corpus of real data results, we present TabDPT-Turbo, a model that provides comparable default performance to TabDPT v1.1 on TabArena-Lite, CC18, and CTR23, at orders of magnitude faster. In our experiments, TabDPT-Turbo is the fastest model overall among leading foundation models. We have released the new model as TabDPT v1.2 at https://github.com/layer6ai-labs/TabDPT-inference.