Schema matching increasingly uses generative language models to rerank retrieved column candidates, although the underlying task is a bounded correspondence decision. We present JevNexus, which combines typed pairwise decisions with schema/instance evidence and invokes listwise refinement only when the evidence disagrees and the fused margin is small. The evaluation covers 561 cases from six benchmark families. JevNexus obtains dataset-macro MRR and Hits@1 of 0.930 and 0.909, compared with 0.926 and 0.903 for Magneto, while reducing mean latency from 123.452 to 15.929 seconds (7.750). Paired analysis finds no statistically significant difference in either MRR or Hits@1. The gate invokes listwise refinement for only 5.665% of source columns and avoids the degradation caused by unconditional refinement. Code and experimental artifacts are available at https://github.com/RazeenLI/JevNexus.
Figures & tables
Figure 1. A schema-matching example with conflicting evidence.
Figure 2. Generative and decision-centric inference over the same localized candidates.
Figure 3. Overview of the JevNexus matching pipeline.
Figure 4. Typed correspondence decisions for two localized candidates.
Dataset
Cases
Source cols.
Source width
Target width
GT pairs
GT/case
ChEMBL
180
3,060
12–23
12–23
2,052
11.4
GDC
10
569
16–179
736
259
25.9
Magellan
7
41
4–9
4–9
41
5.9
OpenData
180
6,876
26–51
26–51
4,572
25.4
TPC-DI
180
2,916
11–22
12–22
1,980
11.0
WikiData
4
60
13–20
13–20
40
10.0
Table 1. Dataset statistics.
Figure 5. Overall matching effectiveness.
ChEMBL
GDC
Magellan
OpenData
TPC-DI
WikiData
Method
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
COMA
0.790
0.715
0.599
0.521
0.481
0.314
1.000
1.000
1.000
0.814
0.781
0.641
0.903
0.881
0.876
0.784
0.748
0.665
COMA++
0.896
0.847
0.747
0.561
0.507
0.366
1.000
1.000
1.000
0.911
0.891
0.514
0.972
0.963
0.870
0.946
0.946
0.873
Distribution
0.549
0.438
0.451
0.000
0.000
0.000
0.686
0.561
0.561
0.563
0.482
0.282
0.790
0.722
0.584
0.808
0.773
0.644
Similarity Flooding
0.793
0.659
0.447
0.544
0.467
0.280
1.000
1.000
1.000
0.650
0.554
0.420
0.877
0.825
0.665
0.766
0.710
0.646
ISResMat
0.838
0.782
0.733
0.448
0.379
0.258
1.000
1.000
1.000
0.686
0.587
0.514
0.868
0.794
0.767
0.946
0.946
0.873
Table 2. Per-dataset matching effectiveness.
Figure 6. Matching effectiveness and inference efficiency.
ChEMBL
GDC
Magellan
OpenData
TPC-DI
WikiData
Variant
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
MRR
H@1
R@GT
w/o Both
0.830
0.766
0.616
0.739
0.642
0.376
0.944
0.895
0.855
0.918
0.882
0.620
0.961
0.931
0.774
0.868
0.817
0.773
w/o Complement
0.869
0.832
0.638
0.797
0.751
0.400
0.990
0.980
0.911
0.948
0.935
0.626
0.941
0.903
0.762
0.889
0.858
0.773
w/o Refinement
0.904
0.873
0.783
0.781
0.703
0.456
1.000
1.000
1.000
0.950
0.935
0.647
0.980
0.963
0.865
0.946
0.946
0.902
Always Refine
0.901
0.865
0.762
0.798
0.730
0.454
1.000
1.000
1.000
0.952
0.939
0.645
0.935
0.889
0.829
0.924
0.902
0.902
JevNexus
0.907
0.878
0.783
0.794
0.725
0.456
1.000
1.000
1.000
0.952
0.939
0.648
0.980
0.964
0.865
0.946
0.946
0.902
Table 3. Component ablation by dataset.
Dataset
Sources
Disagree
Low margin
Calls
Avoided
Activation (%)
ChEMBL
3,060
711
717
111
2,949
3.627
GDC
569
179
189
38
531
6.678
Magellan
41
4
0
0
41
0.000
OpenData
6,876
1,596
1,515
521
6,355
7.577
TPC-DI
2,916
338
786
96
2,820
3.292
WikiData
60
9
7
0
60
0.000
Table 4. Observed online routing by dataset.
Figure 7. Effects of evidence integration and selective refinement.
Figure 8. Hyperparameter sensitivity on the development subset.
Figure 9. Candidate coverage and stage-wise error decomposition.
Schema matching, a critical task for integrating data from diverse sources, seeks to identify correspondences between columns across different schemas. In multi-table holistic schema matching, columns with similar semantic meaning may reside in tables with different contexts due to heterogeneous schema designs, where similarity-based techniques are inadequate. The focus of this paper is exploiting referential context into schema matching by introducing RACT learning and prediction, a self-supervised framework enabling the probabilistic retrieval of candidate tables for source columns to constrain relevant column candidates. Experiments demonstrate that this approach outperforms similarity-based baselines on matching multi-table schemas. In subsequent matching experiments, constraining the column search space via top-t tables improves both average matching precision and completeness by up to +70%.
Leonard Traeger, Enas Khwaileh, Andreas Behrend +1
University of Maryland, Baltimore County, USA · Utrecht University Utrecht, The Netherlands · Technical University of Cologne Cologne, Germany
Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in a multi-database setting, where the system must first locate the target database and then construct a compact, SQL-relevant schema for generation. We propose MDB-Link, a hierarchical schema-linking framework that retrieves question-relevant columns from a global index, aggregates retrieval evidence to shortlist databases, and uses a budget-aware large language model (LLM) for database reranking, table selection, and column grounding. With Qwen2.5-14B, MDB-Link outperforms LinkAlign on MMQA, Spider2-Snow, and BIRD-dev in database localization and column selection while producing schema subsets close in size to the gold schemas. Exact match improves from 16.88 to 51.41 on MMQA, 2.50 to 9.17 on Spider2-Snow, and 12.52 to 38.01 on BIRD-dev. MDB-Link also runs faster than LinkAlign and AutoLink, demonstrating the effectiveness of hierarchical schema reduction for downstream SQL generation.
Beiyu Xu, Zhenyu Wu, Jiaoyan Chen +1
University of Manchester Faculty of Science and Engineering Department of Computer Science
Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task as automated knowledge base construction: (1) prompting an LLM to synthesize a 3NF-compliant relational schema from Wikipedia infobox timelines, (2) populating the schema to obtain a queryable database, and (3) generating and executing SQL queries against it, with QA accuracy serving as an extrinsic evaluation of the constructed knowledge base. In a controlled grid of three schema generators crossed with six query models, the schema source accounts for 79.5% of the exact match (EM) variance against 1.6% for the query model: replacing the schema, and the prompt scaffolding derived from it, shifts EM by 14.7 to 20.0 points, whereas replacing the query model under a fixed schema shifts it by 4.4 to 12.1. From this evidence, we distill three candidate schema-design principles: balanced normalization, semantic naming, and consistent temporal anchoring, framed as correlational hypotheses. Our best configuration (Gemini 2.5 Flash schemas + Gemini-2.0-Flash queries) reaches 80.39 EM, 11.5 points above the strongest reported baseline (68.89 EM); an open-weights configuration reaches 79.52.