cs.LGOct 6, 2026

TICDA: Tabular In-Context Data Attribution

Authors: Yacine Benihaddadene, Milan Bhan, Eliot Dugelay, Mohammed Jawhar, Benjamin Wong, Nicolas Chesneau, Duong Nguyen

Organizations: Ekimetrics · ETH Zurich

Abstract

Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

    Jun 3, 2026Rishit Dagli, Abir Harrasse, Luke Zhang +4Training Data

  2. TAFFY: A Task-Adaptive Tabular Foundation Model with In-Context Diversity

    Oct 6, 2026Zijian Li, Xiangchen Song, Gongxu Luo +7Tabular Foundation ModelsIn-Context Learning