ARCS: Towards Precise Text-to-SQL via Structured Disambiguation
Organizations: Duke University · Megagon Labs
Abstract
As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.
Figures & tables
| Disambiguation | SQL Generation | End-to-end | ||||
| Method | Full Recall (%) | Perfect (%) | EX (%) | EX (%) | EX | Cost ($) |
| Open-source LLMs | ||||||
| gptoss-20b | 4.18 | 0.96 | 27.97 | 7.40 | -20.57 | 0.0003 |
| gptoss-120b | 16.72 | 1.61 | 52.09 | 26.05 | -26.04 | 0.003 |
| qwen3-8b | 6.11 | 0.00 | 11.25 | 5.79 | -5.46 | 0.009 |
| qwen3-235b-a22b-instruct-2507 | 6.43 | 2.57 | 37.30 | 20.26 | -17.04 | 0.02 |
| Method | w/o Taxonomy | w/ Taxonomy | ||
|---|---|---|---|---|
| ARCS EX (%) | Ambrosia EX (%) | ARCS EX (%) | Ambrosia EX (%) | |
| gpt-4.1 | 29.90 | 48.75 | 28.30 | 63.25 |
| o4-mini (medium) | 42.44 | 52.25 | 46.30 | 80.00 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | # Question | # DB | # Table / DB | # Column / DB | # Row / DB | Has Ambiguity? |
|---|---|---|---|---|---|---|
| BIRD (dev) [ 16 ] | 1543 | 11 | 6.82 | 72.55 | 357521 | × |
| Spider 2.0 [ 14 ] | — | — | — | — | — | × |
| Ambrosia [ 23 ] | 1139 | 846 | 4.93 | 19.25 | 22 | ✓ |
| ARCS (ours) | 311 | 6 | 133.83 | 1205.00 | 1969774 | ✓ |
| Measure | Question | Conversational | Structured |
|---|---|---|---|
| Mental Demand | How mentally demanding was the task? | 3.50 | 2.75 |
| Temporal Demand | How hurried or rushed was the pace of the task? | 2.50 | 2.38 |
| Effort | How hard did you have to work to accomplish your level of performance? | 3.50 | 2.63 |
| Frustration | How insecure, discouraged, irritated, stressed, and annoyed were you? | 2.88 | 2.00 |
| Result Confidence | Did you feel confident that the system correctly answered your question? | 5.25 | 5.38 |
| System Transparency | Did the interface make it clear how the system interpreted your input? | 5.13 | 5.75 |