cs.DBSep 30, 2026

TabJoinBench: A Benchmark for Joinable Table Discovery

Authors: Sandipan De, Jin Wang, Vivek Gupta

Organizations: Arizona State University

Abstract

Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery

    Jul 23, 2026Grace Fan, Eden Wu, Majid Daliri +1Data LakesCardinality

  2. JoinGR: Learning to Traverse Join Graphs for Table Retrieval

    Oct 1, 2026Sandipan De, Abhijit Chakraborty, Sambaran Bandyopadhyay +1Graph-Based RetrievalRelational Databases

  3. Discovery-Driven Integration of Disjoint Tables via Text

    Sep 22, 2026Md Ataur Rahman, Dimitris Sacharidis, Oscar Romero +1Data LakesTabular Data