As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.
Figures & tables
Figure 1 : Comparison of conversational disambiguation and structured disambiguation. Structured disambiguation allows users to compare and select interpretations for multiple ambiguity points based on the SQL queries and execution results.
Figure 2 : Illustration of the ARCS taxonomy on a curated database.
Figure 3 : Two task instances from ARCS with annotations of all valid ambiguity points, interpretations, and SQL queries.
Figure 4 : Evaluation settings and metrics for ARCS.
Disambiguation
SQL Generation
End-to-end
Method
Full Recall (%)
Perfect (%)
EX disambiguated (%)
EX e2e (%)
Δ EX
Cost ($)
Open-source LLMs
gptoss-20b
4.18
0.96
27.97
7.40
-20.57
0.0003
gptoss-120b
16.72
1.61
52.09
26.05
-26.04
0.003
qwen3-8b
6.11
0.00
11.25
5.79
-5.46
0.009
qwen3-235b-a22b-instruct-2507
6.43
2.57
37.30
20.26
-17.04
0.02
Table 1 : Main results of structured disambiguation using various LLMs on ARCS. For reasoning models, the level of reasoning effort is indicated in parentheses; ∗ marks the default configuration. Bold indicates the best performance, while underline indicates the second-best. Δ EX denotes the difference between EX e2e and EX disambiguated . Costs are per-task USD price. All open-source LLMs were evaluated using the Fireworks serverless API.
Figure 5 : Comparison of conversational, unstructured, and structured disambiguation. Left: Overall Results. The reasoning effort of o4-mini is medium. Patience is the maximum number of clarification questions the user is willing to answer. User Effort is defined as 0.1 times the user simulator input tokens plus the number of output tokens. Right: End-to-end EX across varying numbers of ambiguity points.
Method
w/o Taxonomy
w/ Taxonomy
ARCS EX (%)
Ambrosia EX (%)
ARCS EX (%)
Ambrosia EX (%)
gpt-4.1
29.90
48.75
28.30
63.25
o4-mini (medium)
42.44
52.25
46.30
80.00
Table 2 : Comparison of end-to-end EX performance on ARCS and Ambrosia, with and without taxonomy as input. Ambrosia is evaluated on a 400-sample subset.
Figure 6 : Error distribution of three representative models on the full ARCS dataset. See Figure 12 for results on the remaining models.
Figure 7 : Ambiguity point recall for each ambiguity type.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
# Question
# DB
# Table / DB
# Column / DB
# Row / DB
Has Ambiguity?
BIRD (dev) [ 16 ]
1543
11
6.82
72.55
357521
×
Spider 2.0 [ 14 ]
—
—
—
—
—
×
Ambrosia [ 23 ]
1139
846
4.93
19.25
22
✓
ARCS (ours)
311
6
133.83
1205.00
1969774
✓
Appendix
Table 3 : Database statistics and comparison with previous benchmarks.
Figure 8 : Benchmark statistics: (left) domain distribution, (middle) distribution of ambiguity points per question, and (right) distribution of SQL queries per question.
Figure 9 : Distribution of ambiguity types across all annotated ambiguity points in ARCS.
Table 4 : Global annotation conventions of ARCS. These instructions are provided to the system during both disambiguation and SQL generation.
Figure 10 : Schema of the retails database.
Figure 11 : Schema of the basketball database.
Table 5 : A sample SQL query at the median of the length distribution in ARCS.
Table 6 : Illustration of the ARCS data format using a sample task entry (Part 1 of 2).
Table 7 : Illustration of the ARCS data format using a sample task entry (Part 2 of 2).
Figure 12 : Error distribution of all evaluated models on the full ARCS dataset.
Table 8 : System prompt for the disambiguation agent used in structured disambiguation. The disambiguation agent has access to the get_schema and final_result tool.
Table 9 : System prompt for the SQL generation agent used in structured disambiguation. The agent has access to the get_schema , get_column_description , search_keywords , run_query and finish tools.
Table 10 : System prompt for the agent used in conversational disambiguation. The agent has access to the ask_user , get_schema , get_column_description , search_keywords , run_query and finish tools.
Table 11 : Prompt used in stage 1 of the user simulator.
Table 12 : Prompt used in stage 2 of the user simulator.
Measure
Question
Conversational
Structured
Mental Demand
How mentally demanding was the task?
3.50
2.75
Temporal Demand
How hurried or rushed was the pace of the task?
2.50
2.38
Effort
How hard did you have to work to accomplish your level of performance?
3.50
2.63
Frustration
How insecure, discouraged, irritated, stressed, and annoyed were you?
2.88
2.00
Result Confidence
Did you feel confident that the system correctly answered your question?
5.25
5.38
System Transparency
Did the interface make it clear how the system interpreted your input?
5.13
5.75
Appendix
Table 13 : User study questionnaire and results. All quantitative measures use a 1–7 scale. For workload measures, lower scores are better; for Result Confidence and System Transparency, higher scores are better.