cs.LGSep 25, 2026

Benchmarking Attention for Tabular Foundation Models

Authors: Maximilian Schambach, Clemens Biehl, Sam Thelin

Organizations: SAP SE, Germany

Abstract

Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention

    Jun 5, 2026Si-Yang Liu, Han-Jia YeTabular Foundation ModelsTabular Prior-Data Fitted Network

  2. Attention Quantization for Tabular Foundation Models

    Sep 14, 2026Jonas M. Kübler, Benjamin Jäger, Klemens Flöge +2Tabular Foundation ModelsKey-Value Cache Quantization

  3. TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction

    Aug 2, 2026Rasa Hosseinzadeh, Alex Labach, Zexin Xue +3Tabular Foundation ModelsIn-Context Learning