cs.CLOct 3, 2026

Understanding Errors in LLM-Based Question Answering over Imperfect Tables

Authors: Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye

Organizations: University of Wisconsin–Madison · University of Auckland, Auckland, New Zealand · Liaoning Provincial People’s Hospital · School of Artificial Intelligence, Jilin University

Abstract

Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language models (LLMs): whether error discovery depends on where errors appear in a table, and whether providing their locations is sufficient for accurate question answering (QA). Using human-reviewed instances from RADAR-T, we conduct controlled studies across three LLMs by varying row order and comparing original, error-marked, and repaired tables. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. During direct inspection, LLMs are more likely to discover all rows containing relevant errors when these rows appear later in the table or are grouped more closely together. Second, providing verified error locations alone is insufficient for accurate QA: with code execution, accuracy on repaired tables exceeds that on error-marked tables by 39.0-59.1 percentage points across the three LLMs. As a practical application of these findings, we combine error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors in a simple workflow, Geometry-Balanced Discovery and Intervention (GBDI). On RADAR-T, GBDI improves QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five LLMs (paired 95% confidence intervals exclude zero for four), at the cost of additional inference. These results highlight the importance of both reliable error discovery and effective error handling in QA over imperfect tables. Code is available at https://github.com/645-t/GBDI-ICLR-2027.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

    Jun 30, 2026Yuqing Yang, Qi Zhu, Zhen Han +5LLM EvaluationLLM Grounding

  2. TabScope: Question-Adaptive Scope Selection for Table Question Answering

    Sep 3, 2026Yuxiang Wang, Junhao Gan, Jianzhong QiTable QA

  3. Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

    May 18, 2026Amritansh Maurya, Navjot Singh, Mohammed Javed +1Table QAStructured Prompting