Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
Organizations: Layer 6 AI
Abstract
Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
Figures & tables
| Dataset | Missing information and construction | Records |
|---|---|---|
| AmbiTab-Bird | Rewrite BIRD tasks with alternative readings; filter by SQL probes and a model judgment. | 664 |
| AmbiTab-EW-Bird | Withhold the BIRD evidence field; retain rows where providing it improves performance for at least one reference model. | 2,537 |
| AmbiTab-BIRD-Interact | Normalize source ambiguity annotations; freeze audited, model-authored responses. | 600 |
| AmbiTab-Ambrosia | Use each annotated reading and its native intent as a separate record. | 2,670 |
| AmbiTab-AmbiQT | Build executable databases for paired SQL readings and assign each reading as a hidden intent. | 1,722 |
| AmbiTab-TACO | Verbalize source SQL while withholding table or value choices. | 187 |
| SR | Precision | Recall | ||||||||
| Dist. | Corpus | Base | RL | Base | RL | Base | RL | |||
| ID | AT-Bird dev | 59.5 | 71.9 | +12.4 | 31.0 | 37.0 | +6.0 | 67.7 | 83.4 | +15.7 |
| AT-EW-Bird dev | 69.0 | 72.2 | +3.2 | 33.6 | 38.0 | +4.4 | 69.8 | 79.3 | +9.5 | |
| OOD | AT-BIRD-Interact | 9.7 | 9.7 | +0.0 | 32.1 | 36.4 | +4.3 | 57.7 | 65.1 | +7.4 |
| AT-Ambrosia | 35.0 | 37.3 | +2.2 | 24.0 | 27.3 | +3.3 | 71.7 | 81.8 | +10.1 | |
| AT-AmbiQT | 56.3 | 59.4 | +3.1 | 19.5 | 21.7 | +2.2 | 58.3 | 65.0 | +6.7 | |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Name | Reading |
| Ambiguous verifiable task | ||
| texts | Specifications, ambiguities, questions, responses, resolutions, and solutions | |
| task specification | The public request, visible to the agent | |
| , | interpretations | is their distribution; is the hidden reading of , drawn once per episode |
| , | ambiguities | The finite set of registered ambiguities of , each named by a text |
| resolution map | is what decides for ambiguity | |
| Corpus | Pool | Records | DBs | With ambiguities | Ambiguities |
|---|---|---|---|---|---|
| AmbiTab-Bird | train | 537 | 42 | 496 | 798 |
| dev | 127 | 11 | 119 | 198 | |
| AmbiTab-EW-Bird | train | 2,091 | 67 | 2,091 | 4,417 |
| dev | 446 | 11 | 446 | 911 | |
| AmbiTab-BIRD-Interact | all | 600 | 22 | 593 | 1,755 |
| AmbiTab-Ambrosia | all | 2,670 | 763 | 2,670 | 2,670 |
| Corpus | Construction | Selection and scope |
|---|---|---|
| AmbiTab-Bird | Construct alternative readings and executable alternative SQL for BIRD tasks, then use Qwen to rewrite the task specification and supply fixed resolutions ( Li et al., 2023 ) . | Require a Qwen3-8B reference-result miss without clarification and a match with all resolutions, followed by a blind Gemma-3-27B necessity judgment; retain zero-ambiguity controls. |
| AmbiTab-EW-Bird | Split BIRD evidence on semicolons, turn each hint into a masked knowledge ambiguity, and withhold it from the visible prompt ( Li et al., 2023 ) . | Keep a row if Qwen3-Coder-30B-A3B or Gemma-3-27B fails without evidence and succeeds with it; evaluate only retained ambiguity-bearing rows, with source train and validation pooled as train and source test as dev. |
| AmbiTab-BIRD-Interact | Retain source task specifications, SQL, PostgreSQL databases, and annotations ( Huo et al., 2026 ) ; normalize duplicate terms, repair one reviewed missing fragment, and freeze an audited catalog of LLM-authored natural-language responses. | Response authoring uses the request, schema, reference SQL, and annotations; the full 600-record pool is evaluated, with local database-disjoint slices also available. |
| AmbiTab-Ambrosia | Convert each reading in the official test partition to a record with its native intent and SQL ( Saparina & Lapata, 2024 ) ; derive terms from scope, attachment, or vagueness annotations and copy the SQLite databases. | Exclude few-shot examples; local database-disjoint train and test slices are selected within each ambiguity type toward a 20% test target, and both derive from the upstream test partition. |
| AmbiTab-AmbiQT | For the column-synonym subset, build a SQLite database containing both alternatives, copying original values to one column and resampling the other with replacement ( Bhaskar et al., 2023 ) . | Reject protected key columns and pairs whose SQL readings fail or return equal results; emit both hidden intents for each of 861 retained source items, with the copied column assigned by a per-unit draw. |
| AmbiTab-TACO | Deterministically verbalize US single-table SELECT queries while withholding table or literal-value choices ( Deng et al., 2026 ) . | Keep executable, nonempty queries with a result-distinct table alternative; 187 of 5,159 source files survive, while value alternatives are not separately required to change the result. |
| Variant | Corpora | Matcher and responder |
|---|---|---|
| Deterministic intent | AmbiTab-BIRD-Interact, AmbiTab-Ambrosia, AmbiTab-AmbiQT, AmbiTab-TACO | Match terms and aliases, preferring quoted targets and then earliest mentions; return the selected frozen response verbatim, with refusals for unmatched questions or ties. |
| Deterministic knowledge | AmbiTab-Bird, AmbiTab-EW-Bird | Rank terms and evidence-derived concepts, prefer unresolved matches, and return the stored resolution; repeats can restate evidence. |
| Scoped LLM | AmbiTab-BIRD-Interact, AmbiTab-Ambrosia, AmbiTab-AmbiQT, AmbiTab-TACO | Classify from routing fields and respond from retained annotations; the responder sees all retained annotations and receives reference SQL when SQL grounding is used. |
| Unfiltered LLM | AmbiTab-Bird, AmbiTab-EW-Bird | Classify and respond with the task context, withheld evidence, and reference SQL available. |
| Setting | Value |
|---|---|
| Asker | Qwen3-8B |
| Solver | Qwen3-Coder-30B-A3B-Instruct, fixed |
| Oracle | Deterministic, Qwen3-14B, or Gemma-3-27B |
| Default asking prompt | Checklist (ledger) |
| Asker and solver decoding | Temperature 0.7; top- (trained askers in Section D.5 : top- ) |
| Qwen3-14B oracle decoding | Temperature 0.7; top- ; top- |
| Strategy | Instructions and history |
|---|---|
| Direct | Ask about unclear parts of the request without an explicit checklist; stop when the request is clear or no different ambiguity remains to be raised. |
| Checklist (chat) | Enumerate potentially ambiguous terms; read prior turns as native chat messages; stop when all terms are resolved or no distinct unasked term remains. |
| Checklist (ledger) | Enumerate potentially ambiguous terms; read the original request and explicit lists of questions with resolved and unavailable responses; continue while an ambiguity remains open, subject to the budget. |
| Setting | Value |
|---|---|
| Initialization / backend | Qwen3-8B / FSDP |
| Optimizer | AdamW, , |
| Learning rate / schedule | Peak ; cosine decay to 0, 5% linear warmup |
| Schedule horizon | 136 optimizer updates, estimated from four decisions per trajectory |
| Weight decay / gradient clipping | 0 / 1.0 |
| LoRA rank / alpha / dropout | 32 / 64 / 0 (adapter scaling ) |
| Policy | SR | Precision | Recall | F1 |
|---|---|---|---|---|
| Question-F1 reward (untrained) | 54.3 | 32.2 | 70.4 | 43.7 |
| Question-F1 reward | 60.6 | 36.8 | 83.3 | 50.6 |
| Outcome reward (deterministic oracle) (untrained) | 55.9 | 31.0 | 68.1 | 42.2 |
| Outcome reward (deterministic oracle) | 57.5 | 33.3 | 73.0 | 45.3 |
| Outcome reward (LLM oracle) (untrained) | 52.8 | 30.6 | 68.1 | 41.9 |
| Outcome reward (LLM oracle) | 57.5 | 31.8 | 70.2 | 43.4 |
| SQL success | |||||||
|---|---|---|---|---|---|---|---|
| Policy | AmbiTab- Bird | AmbiTab- EW-Bird | AmbiTab- BIRD-Interact | AmbiTab- Ambrosia | AmbiTab- AmbiQT | AmbiTab- TACO | Mean |
| Untrained | 59.5 1.4 | 69.0 0.8 | 9.7 0.3 | 35.0 0.2 | 56.3 0.7 | 72.8 1.5 | 50.4 0.4 |
| Question-F1 reward | 71.9 0.5 | 72.2 1.0 | 9.7 0.3 | 37.3 0.2 | 59.4 0.3 | 77.0 1.4 | 54.6 0.2 |
| Outcome reward (deterministic oracle) | 62.7 1.2 | 71.1 1.6 | 9.8 0.1 | 35.5 0.7 | 56.8 0.2 | 74.3 1.9 | 51.7 0.4 |
| Outcome reward (LLM oracle) | 64.3 1.2 | 71.4 1.6 | 9.7 0.6 | 35.9 0.6 | 56.8 0.5 | 73.4 0.8 | 51.9 0.6 |
| Precision | |||||||
| Oracle | Dataset | ||||
|---|---|---|---|---|---|
| Qwen3-14B | AmbiTab-Bird dev | ||||
| AmbiTab-EW-Bird dev | |||||
| AmbiTab-BIRD-Interact | |||||
| AmbiTab-Ambrosia | |||||
| AmbiTab-AmbiQT | |||||
| AmbiTab-TACO |
| Difference, reference SQL shown minus hidden | |||||
|---|---|---|---|---|---|
| Oracle | Dataset | New run | |||
| Qwen3-14B | AmbiTab-Bird dev | SQL hidden | 3.15 [ , 6.56] | 0.52 [ , 3.67] | 0.28 [0.00, 1.12] |
| AmbiTab-EW-Bird dev | SQL hidden | 3.96 [1.42, 6.35] | 2.84 [0.22, 5.61] | 0.00 [ , 0.22] | |
| AmbiTab-Ambrosia | SQL shown | 2.87 [0.88, 4.95] | 4.99 [3.00, 7.02] | 0.25 [ , 1.52] | |
| AmbiTab-AmbiQT | SQL shown | 0.81 [ , 1.72] | 1.24 [0.43, 1.96] | [ , ] | |
| AmbiTab-TACO | SQL shown | [ , 2.50] | [ , 1.96] | [ , 0.00] | |
| AmbiTab-Bird dev | AmbiTab-EW-Bird dev | AmbiTab-Ambrosia | AmbiTab-AmbiQT | AmbiTab-TACO | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Oracle | Asking policy | ||||||||||
| Deterministic | Untrained | 0 | 0 | 0 | 0 | 0 | |||||
| Question-F1 reward | 0 | 0 | 0 | 0 | 0 | ||||||
| Outcome reward (Qwen3-14B oracle) | 0 | 0 | 0 | 0 | 0 | ||||||
| Qwen3-14B | Untrained | ||||||||||
| Question-F1 reward | |||||||||||
| Oracle | Diagnostic | AmbiTab-Bird dev | AmbiTab-EW-Bird dev | AmbiTab-Ambrosia | AmbiTab-AmbiQT | AmbiTab-TACO |
|---|---|---|---|---|---|---|
| Qwen3-14B | 3.94 [ , 10.24] | [ , 5.23] | 2.18 [ , 5.08] | 0.87 [ , 2.19] | 0.71 [ , 4.10] | |
| [ , 0.00] | 0.00 [0.00, 0.00] | [ , ] | 1.18 [ , 3.51] | [ , 0.22] | ||
| Gemma-3-27B | 1.31 [ , 6.82] | 2.77 [ , 6.21] | 0.25 [ , 3.17] | 2.65 [1.32, 4.03] | [ , 3.21] | |
| 1.82 [0.44, 3.81] | 0.07 [ , 0.25] | [ , 5.07] | [ , ] | [ , ] |