LoopICL: Looping a single transformer block to solve tabular tasks
Authors: Amir Rezaei Balef, Katharina Eggensperger
Organizations: TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from computational depth. LoopICL consists of a single block, processing data through two coupled streams: a cell stream capturing per-cell feature representations and a row stream capturing in-context example representations, jointly refined through within-column and cross-column attention. During pre-training, we vary loop counts, allowing the block to be unrolled for a varying number of iterations at test-time and use a learned exit-gate to automatically exit. In its standard setting, LoopICL performs competitively with TabICLv2 on TabArena and TALENT at the same computational cost (FLOPs), while using nearly 90% fewer parameters. Furthermore, its recurrent design enables users to also trade off inference cost and performance, providing a resource-aware TFM.
Figures & tables
Figure 1: Left: LoopICL architecture overview. The blue boxes operate on the cell stream (per-cell representations) ; the red box operates on the row stream (per-row ICL representations) . Right: LoopICL is highly parameter-efficient by reusing the same parameters across recurrent steps.
Figure 2: The number of recurrent iterations L is sampled from a log-normal Poisson distribution during pretraining.
Figure 3: Ablation study. Training stability (gradient norm) and predictive performance (normalized log loss averaged across tasks) of recurrent and non-recurrent model variants.
Figure 4: Residual scaling improves extrapolation. Left: per-iteration norm. log loss for L=12 . Middle: norm. log loss for L∈[1,20] . Right: win rate relative to performance at L=6 , for L≥6 .
Figure 5: Residual scaling ablation. Left: Performance across loop steps for different L and residual scaling schemes. Stars indicate the optimal loop step. Right: Optimal loop steps across datasets, showing dataset-dependent computation and the effect of L on budget allocation.
Figure 6: Latent representation trajectories. PCA of embeddings across loops for binary tasks. Residual scaling with α=\nicefrac1L induces structured and uniform progress toward separating classes.
Figure 7: Improvability (lower is better) measures the relative error gap to the best method, averaged across datasets. Time is training + inference. For TALENT , all models are evaluated on an A100 GPU under identical conditions, enabling a direct runtime comparison. For TabArena , competitor runtimes are taken from the published benchmark and may reflect different hardware; these results should not be used for direct runtime comparisons of LoopICL with other baselines.
Figure 8: Left: Median speedup of LoopICL -EE over the full 16-loop baseline across different λ ; color indicates the fraction of datasets where EE wins. Center / Right: Per-dataset speedup vs. relative error reduction for λ=90 on TabArena and TALENT , compared to the 12-loop baseline (the best-performing fixed-loop configuration). Each point is one dataset; x>1 means EE is faster, y>0 means EE is more accurate. Points in the green quadrant achieve both simultaneously.
Figure 9: Length generalization. Recurrent models demonstrate stronger length generalization compared to the stacked model.
Figure 10: Cross-step embedding similarity for stacked blocks ( left ) and looped models ( right ). Each heatmap entry (i,j) shows the cosine similarity (lower triangle) or linear CKA (upper triangle) between embeddings at loop steps i and j , averaged over all binary-classification datasets.
Figure 11: Cross-step probing AUC for stacked blocks ( left ) and looped models ( right ). Each entry (i,j) shows the ROC-AUC of a logistic regression trained on embeddings at step i and evaluated on embeddings at step j , averaged over all binary-classification datasets.
Figure 12: Left : class separation gap (mean inter-class minus intra-class cosine distance) at each loop step. Right : embedding update magnitude across iterations, measured as the normalised ℓ2 change between consecutive steps ( left panel ) and as the cosine similarity between consecutive embeddings ( right panel ).
Figure 13: With residual scaling L1 we see a correlation between the optimal number of loops and dataset size.
Figure 14: Structure of information in the iterative latent space. We apply PCA to query embeddings pooled across loop steps and samples, and measure the affinity of each principal component (PC) with four functional roles: time-step progression, class separation, prediction confidence, and sample identity.
Figure 15: Elo ratings for classification on TabArena . Higher is better.
Figure 16: Critical difference diagram for classification on TabArena , computed via the autorank framework ( Herbold, 2020 ) using a Wilcoxon signed-rank test with Holm correction ( α=0.05 ). Models connected by a horizontal bar are not significantly different.
Figure 17: Pairwise win rates on TabArena classification (top-20 models). Each cell shows the fraction of datasets where the row model outperforms the column model.
Figure 18: Elo ratings for classification on TALENT . Higher is better.
Figure 19: Critical difference diagram for classification on TALENT , computed via the autorank framework ( Herbold, 2020 ) using a Wilcoxon signed-rank test with Holm correction ( α=0.05 ). Models connected by a horizontal bar are not significantly different.
Figure 20: Pairwise win rates on TALENT classification (top-20 models). Each cell shows the fraction of datasets where the row model outperforms the column model.
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany