Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
Figures & tables
Figure 1: Overview of TabJoinBench . Seed join pairs are generated from three complementary data sources (S1), candidate tables are transformed using composable perturbations while preserving the underlying join relationships (S2), and the resulting query set, heterogeneous data lake, and ground-truth annotations are assembled into the final benchmark (S3).
Subset
Query
Lake
WikiTables (Semantic)
Tables: 293 Avg. Rows: 60.1 Avg. Cols: 8.1
Tables: 11932 Avg. Rows: 55.8 Avg. Cols: 8.8
MMQA (EquiJoin)
Tables: 333 Avg. Rows: 632.0 Avg. Cols: 5.3
Tables: 13106 Avg. Rows: 1270.2 Avg. Cols: 5.1
NYC Open Data (Hybrid)
Tables: 70 Avg. Rows: 1015.2 Avg. Cols: 14.8
Tables: 2638 Avg. Rows: 856.8 Avg. Cols: 10.8
Table 1: Summary statistics of the TabJoinBench benchmark, including the numbers of query and lake tables and their average rows and columns.
Subset
No. of Queries
Relevant Pairs
Relevant / Query
WikiTables
295
11,916
40.4
MMQA
408
13,190
32.3
NYC Open Data
70
2,638
37.7
Table 2: Retrieval characteristics of the TabJoinBench subsets: number of unique queries, total number of relevant candidates, and average number of relevant candidates per query.
Method
P@1
R@1
P@5
R@5
P@15
R@15
P@30
R@30
P@R
(a) WikiTables Subset
LSH Ensemble
0.969
0.038
0.959
0.178
0.950
0.462
0.937
0.516
0.529
PEXESO
0.629
0.026
0.640
0.133
0.630
0.391
0.492
0.488
0.505
D3L
0.199
0.007
0.175
0.026
0.150
0.063
0.125
0.089
0.101
FREYJA
0.660
0.025
0.628
0.119
0.593
0.334
0.398
0.409
0.437
WarpGate
0.220
0.007
0.205
0.034
0.177
0.085
0.140
0.123
0.133
Table 3: Overall retrieval performance on the three TabJoinBench subsets. The best and second-best results within each subset are shown in bold and underlined, respectively.
Method
Struct.
Rep.
Sem.
WikiTables Subset
LSH Ensemble
0.749
0.482
0.534
PEXESO
0.778
0.549
0.471
D3L
0.150
0.109
0.090
FREYJA
0.551
0.417
0.391
WarpGate
0.156
0.150
0.103
Table 4: P@R under each perturbation category (Structural, Representation, Semantic) across TabJoinBench subsets.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Perturbation Type
Relevant Pairs
Relevant / Query
WikiTables (295 Queries)
Structural
3,995
13.54
Representation
3,677
12.46
Semantic
4,244
14.39
Total
11,916
40.39
MMQA (408 Queries)
Structural
4,586
11.24
Representation
4,195
10.28
Appendix
Table 5: Number of relevant candidate pairs and average relevant candidates per query, by perturbation category, across TabJoinBench subsets.
Method
P@1
R@1
P@5
R@5
P@15
R@15
Structural Perturbations
LSH Ensemble
0.947
0.125
0.956
0.621
0.936
0.744
PEXESO
0.784
0.107
0.793
0.541
0.571
0.767
D3L
0.168
0.017
0.166
0.085
0.124
0.152
FREYJA
0.625
0.080
0.609
0.389
0.340
0.545
WarpGate
0.190
0.019
0.172
0.088
0.129
0.173
Appendix
Table 6: Retrieval performance on the perturbed segments of WikiTables Subset.
Method
P@1
R@1
P@5
R@5
P@15
R@15
Structural Perturbations
LSH Ensemble
0.177
0.022
0.174
0.105
0.172
0.225
PEXESO
0.067
0.013
0.062
0.055
0.050
0.098
D3L
0.341
0.044
0.362
0.231
0.289
0.503
FREYJA
0.142
0.017
0.139
0.084
0.125
0.212
WarpGate
0.324
0.041
0.376
0.239
0.300
0.523
Appendix
Table 7: Retrieval performance on the perturbed segments of MMQA Subset.
Method
P@1
R@1
P@5
R@5
P@15
R@15
Structural Perturbations
LSH Ensemble
0.943
0.075
0.943
0.375
0.728
0.725
PEXESO
0.367
0.029
0.36
0.141
0.289
0.325
D3L
0.843
0.067
0.851
0.339
0.719
0.827
FREYJA
0.900
0.071
0.897
0.353
0.620
0.729
WarpGate
0.871
0.069
0.700
0.276
0.586
0.690
Appendix
Table 8: Retrieval performance on the perturbed segments of NYC Open Data Subset.