Understanding Errors in LLM-Based Question Answering over Imperfect Tables
Organizations: University of Wisconsin–Madison · University of Auckland, Auckland, New Zealand · Liaoning Provincial People’s Hospital · School of Artificial Intelligence, Jilin University
Abstract
Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language models (LLMs): whether error discovery depends on where errors appear in a table, and whether providing their locations is sufficient for accurate question answering (QA). Using human-reviewed instances from RADAR-T, we conduct controlled studies across three LLMs by varying row order and comparing original, error-marked, and repaired tables. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. During direct inspection, LLMs are more likely to discover all rows containing relevant errors when these rows appear later in the table or are grouped more closely together. Second, providing verified error locations alone is insufficient for accurate QA: with code execution, accuracy on repaired tables exceeds that on error-marked tables by 39.0-59.1 percentage points across the three LLMs. As a practical application of these findings, we combine error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors in a simple workflow, Geometry-Balanced Discovery and Intervention (GBDI). On RADAR-T, GBDI improves QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five LLMs (paired 95% confidence intervals exclude zero for four), at the cost of additional inference. These results highlight the importance of both reliable error discovery and effective error handling in QA over imperfect tables. Code is available at https://github.com/645-t/GBDI-ICLR-2027.
Figures & tables
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| RADAR task ID and question summary | Source | Reviewed | Position | Dispersion |
|---|---|---|---|---|
| actor-age-gaps Mean age gap for couples with an older male actor. | D12 | M B O F L | M O F L | M O F L |
| board-games-min-players Count games meeting both player-count limits. | D11 | M B O F L | M B L | M B L |
| car-co2 Mean estimated car emissions for 2018–2023. | D06 | M B O F L | F | F |
| daily-activity Mean distance share from moderate-or-higher activity. | D20 | M B O F L | O F | O F |
| employee-years Count older employees with at least five years of tenure. | D25 | M B O F L | — | — |
| england-wales-ethnicity Find the area ranked seventh by Black Caribbean population. | D14 | M B O F L | — | — |
| Source | Dataset |
|---|---|
| D01 | 2021 Green Taxi Trip Data |
| D02 | 2014–15 to 2017–19 NYC Regents Exam Results – Public |
| D03 | Emissions from Industrial Facilities in Queensland – 2004 |
| D04 | Traffic Violations |
| D05 | Tracking data, Subject (a) MC Motion |
| D06 | Fuel Economy Data |
| System | API identifier | Provider |
|---|---|---|
| Qwen Plus | qwen-plus | Bailian |
| MiniMax-M2.5 | MiniMax-M2.5 | Bailian |
| GPT-5 Mini | openai/gpt-5-mini | OpenRouter/OpenAI |
| DeepSeek-V4.1-Flash | deepseek-flash | DeepSeek native API |
| GLM-4.5-Air | glm-4.5-air | Bailian |
| (a) Position | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| Direct | Code Agent | ||||||||
| System | Layout | CDR | Prec. | Rec. | F1 | CDR | Prec. | Rec. | F1 |
| Qwen Plus | q1 | 61.1 | 70.6 | 76.8 | 70.2 | 66.1 | 68.8 | 72.4 | 67.5 |
| q2 | 68.9 | 72.4 | 77.8 | 71.9 | 75.6 | 75.7 | 80.1 | 74.8 | |
| q3 | 66.7 | 71.7 | 76.0 | 70.9 | 76.1 | 76.1 | 82.2 | 76.9 | |
| q4 | 70.6 | 75.5 | 79.5 | 74.7 | 73.9 | 75.4 | 79.2 | 75.1 | |
| Target | Compact | Medium | Wide |
|---|---|---|---|
| 25% | 0.2495 / 0.0081 | 0.2495 / 0.0677 | 0.2495 / 0.1268 |
| 50% | 0.5008 / 0.0081 | 0.5008 / 0.0677 | 0.5008 / 0.1268 |
| 75% | 0.7499 / 0.0081 | 0.7499 / 0.0677 | 0.7499 / 0.1268 |
| System | Dispersion | CDR | Prec. | Rec. | F1 | |
|---|---|---|---|---|---|---|
| Qwen Plus | 25% | Compact | 71.2 | 72.7 | 80.2 | 73.3 |
| Medium | 53.1 | 71.9 | 75.8 | 70.5 | ||
| Wide | 52.0 | 73.0 | 74.0 | 70.0 | ||
| 50% | Compact | 70.1 | 75.9 | 81.0 | 75.7 | |
| Medium | 59.3 | 77.2 | 79.1 | 74.9 | ||
| Wide | 54.2 | 76.9 | 77.6 | 73.8 |
| System | Dispersion | CDR | Prec. | Rec. | F1 | |
|---|---|---|---|---|---|---|
| Qwen Plus | 25% | Compact | 68.9 | 72.4 | 76.0 | 71.3 |
| Medium | 70.6 | 72.4 | 76.7 | 72.0 | ||
| Wide | 68.9 | 73.2 | 76.4 | 71.8 | ||
| 50% | Compact | 71.8 | 75.6 | 78.3 | 74.8 | |
| Medium | 68.4 | 78.1 | 80.5 | 77.2 | ||
| Wide | 72.9 | 76.3 | 80.4 | 75.9 |
| Qwen-Turbo | Qwen3-8B | |||||||
| Layout | CDR | Prec. | Rec. | F1 | CDR | Prec. | Rec. | F1 |
| (a) Position | ||||||||
| q1 | 70.7 | 47.2 | 74.0 | 49.5 | 60.2 | 20.9 | 66.1 | 26.0 |
| q2 | 81.3 | 74.6 | 83.2 | 75.7 | 54.5 | 38.9 | 58.8 | 42.1 |
| q3 | 82.9 | 77.7 | 84.0 | 78.4 | 52.8 | 46.1 | 58.3 | 48.4 |
| q4 | 81.3 | 75.8 | 83.5 | 76.8 | 62.6 | 57.1 | 66.1 | 58.9 |
| System | Views | CDR | Prec. | Rec. | F1 | Mean rows | Empty FP |
|---|---|---|---|---|---|---|---|
| Qwen Plus | Repeated-5 | 58.9 | 71.0 | 77.1 | 68.9 | 9.11 | 61.5 |
| Random-5 | 78.7 | 72.5 | 85.6 | 74.6 | 10.03 | 76.9 | |
| GLM-4.5-Air | Repeated-5 | 61.0 | 60.5 | 74.6 | 62.5 | 9.11 | 84.6 |
| Random-5 | 66.7 | 59.7 | 76.7 | 62.3 | 10.24 | 76.9 | |
| MiniMax-M2.5 | Repeated-5 | 38.3 | 47.1 | 48.9 | 45.4 | 4.98 | 38.5 |
| Random-5 | 58.2 | 65.1 | 69.4 | 64.3 | 7.14 | 23.1 |
| (a) View family guidance | ||||
| Views | Generic | Skill | 95% CI | |
| Repeated-5 | 172 (55.0%) | 190 (60.7%) | ||
| Random-5 | 183 (58.5%) | 203 (64.9%) | ||
| (b) Random-5 vs. Repeated-5 | ||||
| Metric | Repeated-5 | Random-5 | 95% CI | |
| Complete discovery rate | 58.9% | 78.7% | ||
| Generic | Skill | ||||
|---|---|---|---|---|---|
| Input type | Repeated-5 | Random-5 | Repeated-5 | Random-5 | |
| Bad values | 53 | 26 | 28 | 26 | 25 |
| Clean | 53 | 46 | 49 | 47 | 49 |
| Logic | 53 | 18 | 23 | 17 | 28 |
| Formatting | 53 | 29 | 28 | 40 | 38 |
| Missingness | 53 | 25 | 25 | 31 | 30 |
| Guidance | Correct | QA (%) | (pp) | 95% CI (pp) |
|---|---|---|---|---|
| Full Skill | 203/313 | 64.9 | – | – |
| Without recovery rules | 186/313 | 59.4 | ||
| Without formatting rules | 190/313 | 60.7 | ||
| Without execution checks | 197/313 | 62.9 |
| System | Views | Guidance | ||||
|---|---|---|---|---|---|---|
| Qwen Plus | Repeated | Generic | 37 | 46 | 27 | 31 |
| Repeated | Skill | 51 | 32 | 28 | 30 | |
| Random | Generic | 57 | 54 | 16 | 14 | |
| Random | Skill | 72 | 39 | 16 | 14 | |
| GLM-4.5-Air | Repeated | Generic | 32 | 54 | 7 | 48 |
| Repeated | Skill | 38 | 48 | 16 | 39 |
| (a) View budget with union ledger ( ) | ||||||
|---|---|---|---|---|---|---|
| CDR | Prec. | Rec. | F1 | Mean rows | Empty FP | |
| 1 | 45.5 | 71.8 | 71.8 | 68.6 | 6.92 | 64.6 |
| 2 | 65.9 | 72.6 | 80.5 | 72.9 | 8.41 | 70.0 |
| 3 | 73.5 | 72.7 | 83.2 | 73.9 | 9.11 | 73.1 |
| 4 | 76.7 | 72.7 | 84.6 | 74.4 | 9.61 | 75.4 |
| 5 | 78.7 | 72.5 | 85.6 | 74.6 | 10.03 | 76.9 |
| (a) Discovery with fixed prefixes | ||||||
|---|---|---|---|---|---|---|
| CDR | Prec. | Rec. | F1 | Mean rows | Empty FP | |
| 0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.00 | 0.0 |
| 1 | 44.7 | 72.9 | 71.1 | 68.6 | 6.51 | 61.5 |
| 3 | 73.0 | 72.1 | 82.6 | 73.1 | 9.05 | 76.9 |
| 5 | 78.7 | 72.5 | 85.6 | 74.6 | 10.03 | 76.9 |
| Position (60 instances) | Dispersion (59 instances) | |||
|---|---|---|---|---|
| System | Mode | Endpoint | Trend | Wide Compact |
| Qwen Plus | Direct | |||
| Code Agent | ||||
| MiniMax-M2.5 | Direct | |||
| Code Agent | ||||
| GPT-5 Mini | Direct | |||
| System | Mean | 95% CI | |||
|---|---|---|---|---|---|
| Qwen Plus | +13.9 | [+8.4, +20.0] | +19.2 | +16.9 | +5.6 |
| MiniMax-M2.5 | +5.8 | [-0.2, +11.7] | +4.5 | +13.0 | +0.0 |
| GPT-5 Mini | +5.8 | [+0.4, +11.5] | +3.4 | +11.9 | +2.3 |
| System | Mode | Raw | Discovery | Action | Discovery vs. Raw | Action vs. Discovery |
|---|---|---|---|---|---|---|
| Qwen Plus | Direct | 16.2 | 18.8 | 29.2 | ||
| Code Agent | 46.8 | 56.5 | 98.1 | |||
| MiniMax-M2.5 | Direct | 35.7 | 31.8 | 67.5 | ||
| Code Agent | 25.3 | 33.8 | 92.9 | |||
| GPT-5 Mini | Direct | 18.8 | 22.7 | 37.0 | ||
| Code Agent | 47.4 | 58.4 | 97.4 |
| System | CDR | 95% CI |
|---|---|---|
| Qwen Plus | ||
| GLM-4.5-Air | ||
| MiniMax-M2.5 |
| Baseline | Test condition | Gain (pp) | ||||
|---|---|---|---|---|---|---|
| System | Direct | Code Agent | GBDI | 95% CI | ||
| Qwen Plus | 11.8 | 54.0 | 64.9 | +53.0 | +10.9 | |
| GLM-4.5-Air | 4.5 | 41.2 | 49.5 | +45.0 | +8.3 | |
| MiniMax-M2.5 | 44.4 | 32.6 | 51.1 | +6.7 | +18.5 | |
| GPT-5 Mini | 27.2 | 56.9 | 60.7 | +33.5 | +3.8 | |
| DeepSeek-V4.1-Flash | 49.8 | 59.7 | 67.7 | +17.9 | +8.0 | |
| System | Reference | 95% CI | |
|---|---|---|---|
| Qwen Plus | Code Agent | ||
| Repeated + Skill | |||
| MiniMax-M2.5 | Code Agent | ||
| Repeated + Skill | |||
| GPT-5 Mini | Code Agent | ||
| DeepSeek-V4.1-Flash | Code Agent |