cs.AISep 28, 2026

TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora

Authors: Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang, Lihua Yu, Zujie Ren, Gang Chen, Junbo Zhao

Organizations: Zhejiang University · Bank of Hangzhou Co., Ltd. · Zhejiang Lab

Abstract

Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depends on surrounding schema and table context. Heterogeneous schemas, value formats, and serializations further weaken one-shot matching. We present TableSeek, a structure-preserving agentic search framework for heterogeneous table corpora. Instead of ranking tables once, an LLM agent iteratively follows sparse clues, inspects schema-preserving previews, identifies schema- and value-level mismatches, and refines its investigation. TableSeek uses cells and schemas as evidence anchors while retaining complete tables as evidence units, enabling fine-grained localization without losing the context required for interpretation and answerability checking. Without relying on retriever training or a precomputed semantic index, TableSeek produces transparent evidence-seeking trajectories and achieves competitive end-to-end performance against strong retrieval-and-reranking pipelines on heterogeneous table benchmarks. These results suggest that active, structure-preserving evidence seeking is a promising paradigm for open-domain table retrieval.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 20, 2026cs.AI

Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG

In retrieval-augmented generation (RAG), semantic relevance asks whether a source matches a query in meaning, while answerability asks whether it contains sufficient information to answer the query. A Semantic-Answerability Gap (SAG) may arise in retrieval when a retriever can reach semantically relevant sources yet fail to identify those that are uniquely answerable. We uncover this gap using tables as a controlled setting, where shared schemas and entities provide strong semantic signals while localized content and row-column bindings distinguish answerable from non-answerable sources. Using TCR-Bench, a controlled sibling-table benchmark, we find that dense retrievers achieve only 18.2% top-1 target retrieval, reducing QA F1 from 0.755 with the oracle table to 0.330 with retrieved top-5 tables. Controlled diagnostics show that retrievers favor semantic volume over sufficiency, respond weakly to row-column binding disruptions, and struggle to distinguish Targets from Siblings. Explicit answerability assessment substantially improves target identification, while fine-tuning shows that answerability is learnable but difficult to transfer without compromising broad semantic retrieval.
May 1, 2026cs.IR

FollowTable: A Benchmark for Instruction-Following Table Retrieval

Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic similarity. With the growing adoption of LLM-based agentic systems, access to structured data is increasingly instruction-driven, where relevance is conditional on explicit content and schema constraints rather than topical similarity alone. We therefore formalize Instruction-Following Table Retrieval (IFTR), a new task that requires models to jointly satisfy topical relevance and fine-grained instruction constraints. We identify two core challenges in IFTR: (i) sensitivity to content scope, such as inclusion and exclusion constraints, and (ii) awareness of schema-grounded requirements, including column semantics and representation granularity--capabilities largely absent in existing retrievers. To support systematic evaluation, we introduce FollowTable, the first large-scale benchmark for IFTR, constructed via a taxonomy-driven annotation pipeline. We further propose a new metric, termed the Instruction Responsiveness Score, to evaluate whether retrieval rankings consistently adapt to user instructions relative to a topic-only baseline. Our results indicate that existing retrieval models struggle to follow fine-grained instructions over tabular data. In particular, they exhibit systematic biases toward surface-level semantic cues and remain limited in handling schema-grounded constraints, highlighting substantial room for future improvements.
Sep 3, 2026cs.CL

TabScope: Question-Adaptive Scope Selection for Table Question Answering

Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive questions are particularly affected by irrelevant table content, while questions requiring broader evidence may still benefit from full-table reasoning. Based on this observation, we propose a question-adaptive framework that dynamically selects between localized and full-table reasoning. The framework constructs question-specific sub-tables through operation-aware table decomposition and uses the predicted question type to determine the appropriate reasoning mode. We further introduce silver reference sub-tables for evaluating evidence selection and construct SLQA, a benchmark based on real-world long tables. Experiments on WikiTQ and SLQA show that localization is particularly effective for lookup and local reasoning questions, while adaptive selection between localized and full-table reasoning achieves the best overall performance. These results highlight that long-table QA requires deciding not only how to localize, but also when to localize. Our code and datasets will be made available upon publication of the paper.