cs.LGSep 28, 2026

Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer

Authors: Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim

Organizations: Korea University

Abstract

Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks

    Apr 23, 2026Liane Vogel, Kavitha Srinivas, Niharika D'Souza +3Tabular LearningText Embeddings

  2. STRABLE: Benchmarking Tabular Machine Learning with Strings

    May 12, 2026Gioia Blayer, Myung Jun Kim, Félix Lefebvre +8Tabular LearningMle-Bench Lite

  3. TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings

    Jul 27, 2026Ayeen Poostforoushan, Liane Vogel, Carsten BinnigText EmbeddingsMulti-Turn Benchmark