Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depends on surrounding schema and table context. Heterogeneous schemas, value formats, and serializations further weaken one-shot matching. We present TableSeek, a structure-preserving agentic search framework for heterogeneous table corpora. Instead of ranking tables once, an LLM agent iteratively follows sparse clues, inspects schema-preserving previews, identifies schema- and value-level mismatches, and refines its investigation. TableSeek uses cells and schemas as evidence anchors while retaining complete tables as evidence units, enabling fine-grained localization without losing the context required for interpretation and answerability checking. Without relying on retriever training or a precomputed semantic index, TableSeek produces transparent evidence-seeking trajectories and achieves competitive end-to-end performance against strong retrieval-and-reranking pipelines on heterogeneous table benchmarks. These results suggest that active, structure-preserving evidence seeking is a promising paradigm for open-domain table retrieval.
Figures & tables
Figure 1: Semantic relevance = evidence sufficiency. The left table mentions Australia and Capital Region but lacks the actual capital city. The right table contains the exact answer Canberra , but uses Seat of Government instead of Capital , resulting in low lexical overlap with the query.
Figure 2: Overview of TableSeek for question answering over heterogeneous table corpora. Stage 1 scores tables against generated search patterns to form a ranked candidate list. Stage 2 inspects table previews and iteratively refines the patterns until a candidate is selected. Stage 3 checks answerability on the selected table, falling back to the next candidate when evidence is insufficient. Here, Fi denotes a table format, and Ti denotes the i -th ranked table.
Dataset
Split
#Queries
#Tables
TCR-Bench
–
209
637
TARGET-TabFact
valid
12,792
1,696
TARGET-TabFact
valid-1k
1,000
1,696
TARGET-TabFact
test
12,779
1,695
TARGET-TabFact
test-1k
1,000
1,695
Table 1: Statistics of benchmarks and table corpora used in our experiments.
Method
Config
Recall
F1
Dense
E4B
0.172
0.118
Rerank (top-3)
E4B+R4B
0.359
0.269
Rerank (top-5)
E4B+R4B
0.421
0.317
TableSeek (5-step)
Qwen3-30B
0.416
0.284
TableSeek (10-step)
Qwen3-30B
0.507
0.361
Table 2: Main results on TCR-Bench (Mixed format) with Qwen3-30B as the QA/backbone model. Recall is R@1 for Dense/Rerank and R@all for TableSeek. Full results with additional backbones are given in the appendix.
Method
Config
Recall
Acc
Acc-Sup
Dense
Stella
0.519
0.721
0.489
Rerank (top-3)
Stella+R4B
0.623
0.752
0.562
Rerank (top-5)
Stella+R4B
0.656
0.754
0.572
TableSeek (5-step)
Qwen3-30B
0.617
0.754
0.566
TableSeek (10-step)
Qwen3-30B
0.648
0.764
0.589
Table 3: Results on TARGET-TabFact-test-1k with Qwen3-30B as the QA/backbone model. Retrieval is R@1 for Dense/Rerank and R@all for TableSeek.
Metric
Dense
Rerank
TableSeek
top-3
top-5
5step
10step
R@1
0.172
0.359
0.421
–
–
DS@1
0.252
0.478
0.540
–
–
R@all
–
–
–
0.416
0.507
DS@all
–
–
–
0.725
0.803
R@final
–
–
–
0.354
0.431
Table 4: Recall and Discriminative Score on TCR-Bench. Single-pass baselines are evaluated at their single retrieval decision (R@1/DS@1); TableSeek is evaluated both across all QA calls (R@all/DS@all) and at the final call used for answer generation (R@final/DS@final).
Figure 3: F1 across table formats and robustness statistics on TCR-Bench, using Qwen3-30B.
Figure 4: Wall-clock time breakdown by stage (dense encoding, reranking, LLM/QA) across methods on TCR-Bench.
Method
Calls (4B)
Calls (30B)
Dense
1055
1055
Rerank (top-3)
1682
1682
Rerank (top-5)
2100
2100
TableSeek (5-step)
918
921
TableSeek (10-step)
1581
1305
Table 5: Total inference calls on TCR-Bench. Dense/Rerank are independent of the QA backbone.
Strategy
R@all
F1
Avg. Steps
Time (h)
5-step setting
5-step (Full)
0.416
0.284
4.407
2.335
w/o fallback
0.364
0.263
3.919
2.094
w/o refine
0.239
0.173
3.555
1.733
match-only
0.359
0.249
4.435
1.841
match-only QA
0.359
0.065
4.445
1.132
Table 6: Ablation study on refine, fallback, and Structure-Preserving mechanisms on TCR-Bench (Mixed format) using Qwen3-30B.
Figure 5: Top-1 hit rate as a function of refinement round. Iterative refinement continues to improve retrieval quality for the stronger Qwen3-30B model well beyond the first round, whereas the smaller Qwen3-4B model plateaus almost immediately, indicating that the marginal benefit of refinement scales with model capability.
Figure 6: Average action counts per query on TCR-Bench for the default TableSeek configuration, excluding the initial ‘ pattern-gen ‘ step (fixed to 1 per query).
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Max LLM steps per query ( max_steps )
5
Candidate tables shown per refine ( n_candidates )
5
Max tables selectable per refine ( max_select_tables )
1
Table head preview length ( preview_chars )
800
Max matched rows per candidate ( max_matched_rows )
5
Max matched rows per pattern ( max_rows_per_pattern )
Table 8: Refine configurations explored in this work.
Parameter
Value
temperature
0.0
max_tokens
8192
Appendix
Table 9: Decoding parameters used across all experiments.
Hyperparameter
Value
Table source ( table_source )
gold
Tables used for answering ( top_k )
1
Appendix
Table 10: Direct QA hyperparameters.
Hyperparameter
Value
Retrieved tables ( top_k )
5
Similarity
Cosine
Embedding normalization
ℓ2
Appendix
Table 11: Dense retrieval hyperparameters.
Hyperparameter
Value
Dense candidates as input
top-3/top-5
Reranked tables ( top_k )
3/5
Appendix
Table 12: Reranking hyperparameters.
Model
TP
Attn. backend
Ctx. len.
Max pref.
Chunk. pref.
Mem. frac.
Other
Qwen3-30B-A3B-Thinking-2507
4
FlashInfer
65,536
65,536
4,096
0.8
-
Qwen3-4B-Instruct-2507
2
FlashInfer
40,960
40,960
4,096
0.85
-
GLM-Z1-32B-0414
4
FlashInfer
32,768
32,768
4,096
0.8
-
Llama-3.1-8B-Instruct
2
Default
40,960
40,960
4,096
0.8
-
Appendix
Table 13: SGLang serving configuration for different models.
Model
TP
Mem. util.
Max len.
Other
Qwen3.5-9B
4
0.9
40,960
Default
Qwen3.6-35B-A3B
4
0.85
32,768
Default
Qwen3.6-27B
4
0.8
40,960
Default
TableGPT-R1
4
0.85
40,960
Default
Appendix
Table 14: vLLM serving configuration for different models.
Method
Config
Format
Recall
F1 (Qwen3-30B)
F1 (Qwen3-4B)
F1 (LLaMA)
Dense
E4B
CSV
0.158
0.116
0.054
0.079
Dense
E4B
Mixed
0.172
0.118
0.064
0.088
Dense
E4B
HTML
0.134
0.122
0.048
0.056
Dense
E4B
Markdown
0.187
0.134
0.077
0.109
Rerank (top-3)
E4B+R4B
CSV
0.354
0.284
0.102
0.173
Rerank (top-3)
E4B+R4B
Mixed
0.359
0.269
0.127
0.172
Appendix
Table 24: Full backbone comparison on TCR-Bench with Qwen3-30B, Qwen3-4B, and LLaMA. Recall is R@1 for Dense/Rerank and R@all for TableSeek.
Method
Config
Recall
F1
Oracle
GLM-Z1
1.000
0.501
Dense
E4B+GLM-Z1
0.172
0.099
Rerank (top-3)
E4B+R4B+GLM-Z1
0.359
0.197
Rerank (top-5)
E4B+R4B+GLM-Z1
0.421
0.264
TableSeek (5-step-5s1-small)
GLM-Z1
0.258
0.170
TableSeek (10-step-5s1-small)
GLM-Z1
0.306
0.180
Appendix
Table 25: Supplementary backbone comparison on TCR-Bench with additional models beyond the three primary backbones. Recall is R@1 for Dense/Rerank and R@all for TableSeek.
Method
Split
Config
Retrieval
Acc
Acc-Sup
Acc-Ref
Dense
test
Qwen3-EB+Qwen3-4B
0.495
0.628
0.407
0.851
Dense
test-1k
Qwen3-EB+Qwen3-4B
0.498
0.610
0.405
0.823
Dense
test
Stella+Qwen3-4B
0.520
0.637
0.435
0.840
Dense
test-1k
Stella+Qwen3-4B
0.519
0.627
0.446
0.815
Dense
test
Qwen3-EB+LLaMA
0.495
0.611
0.308
0.918
Dense
test-1k
Qwen3-EB+LLaMA
0.498
0.598
0.297
0.910
Appendix
Table 26: Results on TARGET-TabFact-test and TARGET-TabFact-test-1k. Retrieval is R@1 for Dense/Rerank and R@all for TableSeek.
Method
Split
Config
Retrieval
Acc
Acc-Sup
Acc-Ref
Dense
valid
Qwen3-EB+Qwen3-4B
0.492
0.633
0.415
0.857
Dense
valid-1k
Qwen3-EB+Qwen3-4B
0.515
0.610
0.395
0.814
Dense
valid
Stella+Qwen3-4B
0.516
0.639
0.434
0.849
Dense
valid-1k
Stella+Qwen3-4B
0.523
0.629
0.418
0.830
Dense
valid
Qwen3-EB+LLaMA
0.492
0.610
0.315
0.914
Dense
valid-1k
Qwen3-EB+LLaMA
0.515
0.610
0.303
0.902
Appendix
Table 27: Results on TARGET-TabFact-valid and TARGET-TabFact-valid-1k. Retrieval is R@1 for Dense/Rerank and R@all for TableSeek.
Comparison
Backbone LLM
F1 A
F1 B
Δ
95% CI
p
TableSeek (10-step) vs Rerank (top-5)
Qwen3.6-27B
0.592
0.359
+0.233
[0.167, 0.302]
<0.001
TableSeek (5-step) vs Rerank (top-3)
Qwen3.6-27B
0.545
0.323
+0.222
[0.152, 0.294]
<0.001
TableSeek (10-step) vs Rerank (top-5)
Qwen3-30B
0.361
0.317
+0.045
[-0.028, 0.118]
0.229
TableSeek (5-step) vs Rerank (top-3)
Qwen3-30B
0.284
0.259
+0.024
[-0.044, 0.095]
0.482
TableSeek (10-step) vs TableSeek (5-step)
Qwen3-30B
0.361
0.284
+0.077
[0.028, 0.128]
0.002
Appendix
Table 28: Paired significance analysis on TCR-Bench using per-instance QA F1 scores. Bootstrap uses 10,000 resamples.
Model
CSV
Markdown
HTML
Mixed
Qwen3-4B
0.239
0.301
0.266
0.270
LLaMA
0.491
0.485
0.417
0.465
Qwen3-30B
0.717
0.748
0.719
0.734
Appendix
Table 29: Oracle QA results (F1) on TCR-Bench under different table formats.
Model
Split
Acc
Acc-Sup
Acc-Ref
Qwen3-4B
valid
0.732
0.690
0.775
Qwen3-4B
valid-1k
0.723
0.683
0.764
Qwen3-4B
test
0.695
0.654
0.737
Qwen3-4B
test-1k
0.695
0.654
0.737
LLaMA
valid
0.699
0.528
0.875
LLaMA
valid-1k
0.684
0.488
0.871
Appendix
Table 30: Oracle QA results on TabFact.
Method
Model(s)
Format
R@1
R@2
R@3
Dense
Qwen3-Embedding-4B
CSV
0.158
0.311
0.464
Dense
Qwen3-Embedding-4B
Mixed
0.172
0.321
0.469
Dense
Qwen3-Embedding-4B
HTML
0.134
0.287
0.474
Dense
Qwen3-Embedding-4B
Markdown
0.187
0.340
0.512
Dense
jina-embeddings-v4
CSV
0.129
0.282
0.368
Dense
jina-embeddings-v4
Mixed
0.163
0.306
0.426
Appendix
Table 31: Retrieval-only results on TCR-Bench.
Method
Model(s)
Split
R@1
R@2
R@3
Dense
Qwen3-Embedding-4B
test
0.495
0.596
0.649
Dense
Qwen3-Embedding-4B
test-1k
0.498
0.589
0.646
Dense
jina-embeddings-v4
test
0.439
0.528
0.579
Dense
jina-embeddings-v4
test-1k
0.450
0.542
0.593
Dense
stella_en_1.5B_v5
test
0.520
0.615
0.664
Dense
stella_en_1.5B_v5
test-1k
0.519
0.614
0.668
Appendix
Table 32: Retrieval-only results on TARGET-TabFact-test and TARGET-TabFact-test-1k.
Method
Model(s)
Split
R@1
R@2
R@3
Dense
Qwen3-Embedding-4B
valid
0.492
0.592
0.647
Dense
Qwen3-Embedding-4B
valid-1k
0.515
0.606
0.641
Dense
jina-embeddings-v4
valid
0.428
0.519
0.570
Dense
jina-embeddings-v4
valid-1k
0.432
0.522
0.579
Dense
stella_en_1.5B_v5
valid
0.516
0.618
0.668
Dense
stella_en_1.5B_v5
valid-1k
0.523
0.611
0.661
Appendix
Table 33: Retrieval-only results on TARGET-TabFact-valid and TARGET-TabFact-valid-1k.
Method
Dense Calls
Rerank Calls
LLM Calls
Total Calls
Dense Time
Rerank Time
LLM Time
Total Time
Dense (Qwen3-4B)
846
0
209
1055
0.940
0
0.125
1.065
Rerank-3 (Qwen3-4B)
846
627
209
1682
0.940
0.964
0.119
2.023
Rerank-5 (Qwen3-4B)
846
1045
209
2100
0.940
1.597
0.122
2.659
Dense (Qwen3-30B)
846
0
209
1055
0.940
0
0.765
1.705
Rerank-3 (Qwen3-30B)
846
627
209
1682
0.940
0.964
0.711
2.615
Rerank-5 (Qwen3-30B)
846
1045
209
2100
0.940
1.597
0.704
3.241
Appendix
Table 34: Number of inference calls and wall-clock runtime (hours) on TCR-Bench.
Strategy
R@all
F1
Avg. Steps
Time (h)
5-step setting
Full
0.416
0.284
4.407
2.335
w/o refine
0.239
0.173
3.555
1.733
w/o regenerate
0.254
0.190
4.148
1.921
w/o select
0.402
0.305
4.909
3.031
10-step setting
Appendix
Table 35: Fine-grained ablation of the refine mechanism, decomposed into regenerate and select, on TCR-Bench (Mixed format) using Qwen3-30B.
Round
Qwen3-30B
Qwen3-4B
Qwen3-30B (TabFact)
GLM-Z1
Qwen3.5-9B
5-step
10-step
5-step
10-step
5-step
10-step
10-step
10-step
0
0.206
0.201
0.153
0.153
0.441
0.444
0.062
0.306
1
0.301
0.301
0.158
0.158
0.488
0.486
0.206
0.349
2
0.311
0.330
0.158
0.158
0.498
0.493
0.206
0.359
3
0.330
0.340
0.163
0.163
0.505
0.493
0.210
0.359
4
–
0.364
–
0.158
–
0.506
0.220
0.364
Appendix
Table 36: Top-1 hit rate by refinement round, for two independent runs (5-step and 10-step budgets) of each model configuration.
Strategy
R@all
F1
Eff
Avg. Steps
Time (h)
5step-5s1
0.416
0.284
0.683
4.407
2.335
5step-5s3
0.445
0.307
0.690
4.407
2.609
5step-10s1
0.397
0.294
0.741
4.383
2.807
5step-10s3
0.431
0.297
0.689
4.407
3.129
5step-10s5
0.469
0.316
0.674
4.354
3.259
10step-5s1
0.507
0.361
0.712
6.244
3.466
Appendix
Table 37: Effect of refine strategy on TCR-Bench (Mixed format), using Qwen3-30B. Eff = F1/R@all.
Strategy
R@all
F1
Eff
Avg. Steps
Time (h)
5step-r3
0.378
0.269
0.712
4.431
2.172
5step-r5
0.416
0.291
0.700
4.459
2.580
5step-r10
0.416
0.307
0.738
4.349
2.850
10step-r3
0.498
0.375
0.753
6.196
3.174
10step-r5
0.550
0.384
0.698
6.067
3.535
10step-r10
0.517
0.374
0.723
6.038
4.069
Appendix
Table 38: Effect of the rerank-based refine strategy (r n ) on TCR-Bench (Mixed format), using Qwen3-30B. Eff = F1/R@all.
Format
10step-5s1
10step-5s3
10step-r5
CSV
0.346
0.367
0.362
Mixed
0.361
0.352
0.384
HTML
0.364
0.356
0.352
Markdown
0.364
0.389
0.346
Appendix
Table 39: F1 of selected 10-step refine strategies across table formats on TCR-Bench, using Qwen3-30B.
Method
Dense
Rerank (top-3)
Rerank (top-5)
TableSeek (5step)
TableSeek (10step)
CSV
0.116
0.284
0.356
0.283
0.346
Mixed
0.118
0.269
0.317
0.284
0.361
HTML
0.122
0.279
0.353
0.292
0.364
Markdown
0.134
0.314
0.362
0.318
0.364
Mean
0.123
0.287
0.347
0.294
0.359
Range
0.018
0.045
0.045
0.035
0.018
Appendix
Table 40: F1 across table formats and robustness statistics on TCR-Bench, using Qwen3-30B.
TableSeek (10-step)
Rerank (top-5)
Format
Run1
Run2
Run3
Avg
Run1
Run2
Run3
Avg
CSV
0.346
0.340
0.340
0.342
0.356
0.351
0.352
0.353
Mixed
0.361
0.360
0.363
0.362
0.317
0.312
0.317
0.315
HTML
0.364
0.364
0.364
0.364
0.353
0.353
0.353
0.353
Markdown
0.364
0.358
0.375
0.366
0.362
0.362
0.362
0.362
Mean
–
–
–
0.358
–
–
–
0.346
Appendix
Table 41: Multi-run robustness analysis for Rerank (top-5) and TableSeek (10-step). Robustness statistics are computed from the format-wise average F1 scores across three independent runs.
Method
10-step-5s3
10-step-r5
CSV
0.367
0.362
Mixed
0.352
0.384
HTML
0.356
0.352
Markdown
0.389
0.346
Mean
0.366
0.361
Range
0.037
0.038
Appendix
Table 42: Robustness analysis of alternative refine strategies.
Configuration
R@All
F1
Avg. Steps
Time (h)
10-step setting
full
0.512
0.371
6.153
10.079
big
0.502
0.370
6.139
3.498
medium
0.493
0.372
6.340
3.606
small
0.507
0.361
6.244
3.466
5-step setting
Appendix
Table 43: Effect of retrieved evidence size on TCR-Bench (Mixed format) using Qwen3-30B. Larger evidence windows generally improve retrieval performance but incur substantially higher runtime costs.
Method
Config
Recall T
F1 T
Oracle
Qwen3-30B
1.000
0.269
Dense
E4B+Qwen3-30B
0.081
0.048
Rerank (top-3)
E4B+R0.6B+Qwen3-30B
0.187
0.076
Rerank (top-5)
E4B+R0.6B+Qwen3-30B
0.239
0.100
TableSeek (5-step-5s1-small)
Qwen3-30B
0.263
0.110
TableSeek (10-step-5s1-small)
Qwen3-30B
0.335
0.139
Appendix
Table 44: Results on the transposed-table version of TCR-Bench.
In retrieval-augmented generation (RAG), semantic relevance asks whether a source matches a query in meaning, while answerability asks whether it contains sufficient information to answer the query. A Semantic-Answerability Gap (SAG) may arise in retrieval when a retriever can reach semantically relevant sources yet fail to identify those that are uniquely answerable. We uncover this gap using tables as a controlled setting, where shared schemas and entities provide strong semantic signals while localized content and row-column bindings distinguish answerable from non-answerable sources. Using TCR-Bench, a controlled sibling-table benchmark, we find that dense retrievers achieve only 18.2% top-1 target retrieval, reducing QA F1 from 0.755 with the oracle table to 0.330 with retrieved top-5 tables. Controlled diagnostics show that retrievers favor semantic volume over sufficiency, respond weakly to row-column binding disruptions, and struggle to distinguish Targets from Siblings. Explicit answerability assessment substantially improves target identification, while fine-tuning shows that answerability is learnable but difficult to transfer without compromising broad semantic retrieval.
Jiaming Tian, Liyao Li, Wentao Ye +6
Zhejiang University · Bank of Hangzhou Co., Ltd. · Zhejiang Lab
Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic similarity. With the growing adoption of LLM-based agentic systems, access to structured data is increasingly instruction-driven, where relevance is conditional on explicit content and schema constraints rather than topical similarity alone. We therefore formalize Instruction-Following Table Retrieval (IFTR), a new task that requires models to jointly satisfy topical relevance and fine-grained instruction constraints. We identify two core challenges in IFTR: (i) sensitivity to content scope, such as inclusion and exclusion constraints, and (ii) awareness of schema-grounded requirements, including column semantics and representation granularity--capabilities largely absent in existing retrievers. To support systematic evaluation, we introduce FollowTable, the first large-scale benchmark for IFTR, constructed via a taxonomy-driven annotation pipeline. We further propose a new metric, termed the Instruction Responsiveness Score, to evaluate whether retrieval rankings consistently adapt to user instructions relative to a topic-only baseline. Our results indicate that existing retrieval models struggle to follow fine-grained instructions over tabular data. In particular, they exhibit systematic biases toward surface-level semantic cues and remain limited in handling schema-grounded constraints, highlighting substantial room for future improvements.
Rihui Jin, Yuchen Lu, Ting Zhang +7
Southeast University Nanjing, China Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education · Huawei Noah’s Ark Lab Shenzhen, China
Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affected by irrelevant table content, while questions requiring broader evidence may still benefit from full-table reasoning. Based on this observation, we propose a question-adaptive framework that dynamically selects between localized and full-table reasoning. The framework constructs question-specific sub-tables through operation-aware table decomposition and uses the predicted question type to determine the appropriate reasoning mode. We further introduce silver reference sub-tables for evaluating evidence selection and construct SLQA, a benchmark based on real-world long tables. Experiments on WikiTQ and SLQA show that localization is particularly effective for lookup and local reasoning questions, while adaptive selection between localized and full-table reasoning achieves the best overall performance. These results highlight that long-table QA requires deciding not only how to localize, but also when to localize. Our code and datasets will be made available upon publication of the paper.