LoopICL: Looping a single transformer block to solve tabular tasks
Authors: Amir Rezaei Balef, Katharina Eggensperger
Organizations: TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from computational depth. LoopICL consists of a single block, processing data through two coupled streams: a cell stream capturing per-cell feature representations and a row stream capturing in-context example representations, jointly refined through within-column and cross-column attention. During pre-training, we vary loop counts, allowing the block to be unrolled for a varying number of iterations at test-time and use a learned exit-gate to automatically exit. In its standard setting, LoopICL performs competitively with TabICLv2 on TabArena and TALENT at the same computational cost (FLOPs), while using nearly 90% fewer parameters. Furthermore, its recurrent design enables users to also trade off inference cost and performance, providing a resource-aware TFM.
Figures & tables
Figure 1: Left: LoopICL architecture overview. The blue boxes operate on the cell stream (per-cell representations) ; the red box operates on the row stream (per-row ICL representations) . Right: LoopICL is highly parameter-efficient by reusing the same parameters across recurrent steps.
Figure 2: The number of recurrent iterations L is sampled from a log-normal Poisson distribution during pretraining.
Figure 3: Ablation study. Training stability (gradient norm) and predictive performance (normalized log loss averaged across tasks) of recurrent and non-recurrent model variants.
Figure 4: Residual scaling improves extrapolation. Left: per-iteration norm. log loss for L=12 . Middle: norm. log loss for L∈[1,20] . Right: win rate relative to performance at L=6 , for L≥6 .
Figure 5: Residual scaling ablation. Left: Performance across loop steps for different L and residual scaling schemes. Stars indicate the optimal loop step. Right: Optimal loop steps across datasets, showing dataset-dependent computation and the effect of L on budget allocation.
Figure 6: Latent representation trajectories. PCA of embeddings across loops for binary tasks. Residual scaling with α=\nicefrac1L induces structured and uniform progress toward separating classes.
Figure 7: Improvability (lower is better) measures the relative error gap to the best method, averaged across datasets. Time is training + inference. For TALENT , all models are evaluated on an A100 GPU under identical conditions, enabling a direct runtime comparison. For TabArena , competitor runtimes are taken from the published benchmark and may reflect different hardware; these results should not be used for direct runtime comparisons of LoopICL with other baselines.
Figure 8: Left: Median speedup of LoopICL -EE over the full 16-loop baseline across different λ ; color indicates the fraction of datasets where EE wins. Center / Right: Per-dataset speedup vs. relative error reduction for λ=90 on TabArena and TALENT , compared to the 12-loop baseline (the best-performing fixed-loop configuration). Each point is one dataset; x>1 means EE is faster, y>0 means EE is more accurate. Points in the green quadrant achieve both simultaneously.
Figure 9: Length generalization. Recurrent models demonstrate stronger length generalization compared to the stacked model.
Figure 10: Cross-step embedding similarity for stacked blocks ( left ) and looped models ( right ). Each heatmap entry (i,j) shows the cosine similarity (lower triangle) or linear CKA (upper triangle) between embeddings at loop steps i and j , averaged over all binary-classification datasets.
Figure 11: Cross-step probing AUC for stacked blocks ( left ) and looped models ( right ). Each entry (i,j) shows the ROC-AUC of a logistic regression trained on embeddings at step i and evaluated on embeddings at step j , averaged over all binary-classification datasets.
Figure 12: Left : class separation gap (mean inter-class minus intra-class cosine distance) at each loop step. Right : embedding update magnitude across iterations, measured as the normalised ℓ2 change between consecutive steps ( left panel ) and as the cosine similarity between consecutive embeddings ( right panel ).
Figure 13: With residual scaling L1 we see a correlation between the optimal number of loops and dataset size.
Figure 14: Structure of information in the iterative latent space. We apply PCA to query embeddings pooled across loop steps and samples, and measure the affinity of each principal component (PC) with four functional roles: time-step progression, class separation, prediction confidence, and sample identity.
Figure 15: Elo ratings for classification on TabArena . Higher is better.
Figure 16: Critical difference diagram for classification on TabArena , computed via the autorank framework ( Herbold, 2020 ) using a Wilcoxon signed-rank test with Holm correction ( α=0.05 ). Models connected by a horizontal bar are not significantly different.
Figure 17: Pairwise win rates on TabArena classification (top-20 models). Each cell shows the fraction of datasets where the row model outperforms the column model.
Figure 18: Elo ratings for classification on TALENT . Higher is better.
Figure 19: Critical difference diagram for classification on TALENT , computed via the autorank framework ( Herbold, 2020 ) using a Wilcoxon signed-rank test with Holm correction ( α=0.05 ). Models connected by a horizontal bar are not significantly different.
Figure 20: Pairwise win rates on TALENT classification (top-20 models). Each cell shows the fraction of datasets where the row model outperforms the column model.
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific architectures reduces model size and computational demands but also sacrifices in-context adaptability. Here we introduce TACTICL, an automated task-aware compression framework for tabular in-context learning models that jointly prunes transformer layers and replaces them with lightweight adapters trained on downstream tasks, thus blending in-context with in-weight learning. We study TACTICL on 47 benchmark datasets and show that we can substitute up to 85% of layers without substantial performance drop on a given downstream task. We further show that TACTICL maintains robustness to data shifts, leaving its in-context ability intact. Overall, TACTICL provides a robust framework for exploiting the depth-wise redundancy of tabular foundation models by combining task-specific adaptation and structured compression. We provide the code at: https://github.com/Hebog/tfm_compression
Tabular foundation models, such as TabPFNv2 and TabICL, have recently dethroned gradient-boosted trees at the top of predictive benchmarks, demonstrating the value of in-context learning for tabular data. We introduce TabICLv2, a new state-of-the-art foundation model for regression and classification built on three pillars: (1) a novel synthetic data generation engine designed for high pretraining diversity; (2) various architectural innovations, including a new scalable softmax in attention improving generalization to larger datasets without prohibitive long-sequence pretraining; and (3) optimized pretraining protocols, notably replacing AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 without any tuning surpasses the performance of the current state of the art, RealTabPFN-2.5 (hyperparameter-tuned, ensembled, and fine-tuned on real data). With only moderate pretraining compute, TabICLv2 generalizes effectively to million-scale datasets under 50 GB GPU memory while being markedly faster than RealTabPFN-2.5. We provide extensive ablation studies to quantify these contributions and foster open research by releasing code for inference, pretraining, and synthetic data generation at https://github.com/soda-inria/tabicl.
Jingang Qu, David Holzmüller, Gaël Varoquaux +1
SODA Team, INRIA Saclay, Palaiseau, France · Probabl, France