Why, Where, How: Taxonomy-guided Error Grounding for Code Repair in NL2SQL
Organizations: POSTECH · KAIST
Abstract
SQL queries that large language models write from natural language questions can execute successfully yet produce incorrect results, so execution alone does not reveal what to fix. An error taxonomy says why the query is wrong, but not where to look or how to change it. Existing methods can guide SQL correction through feedback, error reports, or generated plans alongside an unmasked query. We introduce TEG(Taxonomy-guided Error Grounding), which turns a supplied diagnosis into a structured correction input for natural language-to-SQL (NL2SQL) correction. Type-specific rules map each error type to construct classes to reconsider and an edit operation to request. TEG masks the selected constructs in the query when applicable and states that operation in an edit instruction. TEG generates candidate corrections from this input, uses execution feedback to guide candidate selection, and repeats the process one annotation at a time for queries with several errors. On NL2SQL-BUGs, TEG reaches 47.3 single-error execution accuracy and 37.0 overall with Qwen2.5-7B-Instruct. Across the model sizes and thinking modes evaluated in the main comparison, TEG outperforms all evaluated baselines on single-error queries, even when the baselines receive the same error-type annotations. With predicted types, TEG stays above direct LLM correction and ErrorLLM on single-error queries.
Figures & tables
| Qwen2.5-Instruct | Qwen3-8B | |||||
| Method | Error types | 7B | 14B | 32B | Thinking OFF | Thinking ON |
| Base | ✗ | 30.0 | 35.5 | 41.8 | 31.8 | 43.6 |
| Reasoning and self-correction | ||||||
| CoT ( Kojima et al., 2022 ) | ✗ | 35.5 | 37.3 | 45.5 | 33.6 | 48.2 |
| Self-Debug (Simple) ( Chen et al., 2024b ) | ✗ | 33.6 | 38.2 | 48.2 | 38.2 | 50.9 |
| Self-Debug (Expl) ( Chen et al., 2024b ) | ✗ | 29.1 | 35.5 | 44.5 | 44.5 | 47.3 |
| Method | Error types | 1 error | 2 errors | 3+ errors | Overall |
| Base | ✗ | 30.0 | 25.9 | 14.1 | 23.7 |
| Reasoning and self-correction | |||||
| CoT ( Kojima et al., 2022 ) | ✗ | 35.5 | 21.5 | 21.8 | 24.4 |
| Self-Debug (Simple) ( Chen et al., 2024b ) | ✗ | 33.6 | 20.9 | 16.9 | 22.4 |
| Self-Debug (Expl) ( Chen et al., 2024b ) | ✗ | 29.1 | 24.6 | 18.3 | 23.9 |
| SQL-specific refinement | |||||
| Qwen2.5-Instruct | Qwen3-8B | Opus 4.8 | ||||
| Method | 7B | 14B | 32B | Thinking OFF | Thinking ON | |
| Base | 30.0 | 35.5 | 41.8 | 31.8 | 43.6 | 63.6 |
| ErrorLLM | 32.7 | 45.5 | 47.3 | 38.2 | 44.5 | 64.5 |
| TEG (ours) | 45.5 | 48.2 | 52.7 | 45.5 | 53.6 | 69.1 |
| Method | 1 error | 2 errors | 3+ errors | Overall |
| TEG (ours) | 47.3 | 36.7 | 29.6 | 37.0 |
| w/o error annotation | 38.2 | 34.7 | 21.1 | 31.9 |
| w/o masking | 40.9 | 34.3 | 27.5 | 33.9 |
| w/o DB execution | 40.9 | 27.3 | 19.0 | 27.9 |
| w/o candidate selection | 47.3 | 25.9 | 20.4 | 28.8 |
| w/o sequential correction | 47.3 | 32.3 | 23.9 | 33.2 |
| Method | Error types | Attribute | Value | Condition | Table | Others | Overall |
| Base | ✗ | 24.3 | 41.4 | 29.4 | 33.3 | 16.7 | 30.0 |
| CoT | ✗ | 35.1 | 37.9 | 52.9 | 26.7 | 16.7 | 35.5 |
| Self-Debug (Simple) | ✗ | 24.3 | 41.4 | 35.3 | 53.3 | 16.7 | 33.6 |
| Self-Debug (Expl) | ✗ | 24.3 | 37.9 | 35.3 | 33.3 | 0 8.3 | 29.1 |
| DIN-SQL | ✗ | 10.8 | 51.7 | 23.5 | 13.3 | 16.7 | 24.5 |
| MAC-SQL | ✗ | 10.8 | 27.6 | 23.5 | 0 6.7 | 25.0 | 18.2 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Sub-type | Edit group |
| Attribute-Related | Attribute Mismatch | M |
| Attribute Redundancy | R | |
| Attribute Missing | Mi | |
| Table-Related | Table Mismatch | M |
| Table Redundancy | R | |
| Table Missing | Mi |
| Category | Sub-type | Target AST Nodes |
| Attribute | Mismatch, Redundancy | Column (string-level replace) |
| Missing | – | |
| Table | Mismatch, Redundancy | Table |
| Join Cond./Type Mismatch | Join | |
| Missing | – | |
| Value | Mismatch, Data Format | Literal , String , Int64 |
| Distribution | Count | % |
| Errors per Query (549 queries) | ||
| 1 error | 110 | 20.0 |
| 2 errors | 297 | 54.1 |
| 3 errors | 114 | 20.8 |
| 4 errors | 26 | 4.7 |
| 5 errors | 2 | 0.4 |
| Example 1 : Attribute Mismatch (single error) db: california_schools | |
| Question | What is the unabbreviated mailing street address of the school with the highest FRPM count for K-12 students? |
| Issued SQL | SELECT schools. streetabr FROM frpm INNER JOIN schools ON frpm.cdscode = schools.cdscode ORDER BY frpm.‘frpm count (k-12)‘ DESC LIMIT 1 |
| Reference SQL | SELECT T2. MailStreet FROM frpm AS T1 INNER JOIN schools AS T2 ON T1.CDSCode = T2.CDSCode ORDER BY T1.‘FRPM Count (K-12)‘ DESC LIMIT 1 |
| Example 2 : Value Mismatch (single error) db: codebase_community | |
| Question | Which user has the website URL listed at ’http://stackoverflow.com’ |
| Issued SQL | SELECT displayname FROM users WHERE websiteurl = ’http://stackoverflow.com/u/1114’ |
| Method | Error types | Attribute | Value | Condition | Table | Others | Overall |
| Qwen2.5-7B | |||||||
| Base | ✗ | 24.3 | 41.4 | 29.4 | 33.3 | 16.7 | 30.0 |
| CoT | ✗ | 35.1 | 37.9 | 52.9 | 26.7 | 16.7 | 35.5 |
| Self-Debug (Simple) | ✗ | 24.3 | 41.4 | 35.3 | 53.3 | 16.7 | 33.6 |
| Self-Debug (Expl) | ✗ | 24.3 | 37.9 | 35.3 | 33.3 | 8.3 | 29.1 |
| DIN-SQL | ✗ | 10.8 | 51.7 | 23.5 | 13.3 | 16.7 | 24.5 |
| Method | Error types | Attribute | Value | Condition | Table | Others | Overall |
| Qwen3-8B (Thinking OFF) | |||||||
| Base | ✗ | 27.0 | 34.5 | 29.4 | 46.7 | 25.0 | 31.8 |
| CoT | ✗ | 18.9 | 37.9 | 41.2 | 66.7 | 16.7 | 33.6 |
| Self-Debug (Simple) | ✗ | 40.5 | 44.8 | 29.4 | 33.3 | 33.3 | 38.2 |
| Self-Debug (Expl) | ✗ | 43.2 | 44.8 | 35.3 | 66.7 | 33.3 | 44.5 |
| DIN-SQL | ✗ | 13.5 | 41.4 | 17.6 | 13.3 | 16.7 | 21.8 |
| Method | Attribute | Value | Condition | Table | Others | Overall |
| Base | 19.6 | 19.4 | 15.7 | 24.2 | 22.1 | 22.1 |
| Reasoning and self-correction | ||||||
| CoT ( Kojima et al., 2022 ) | 20.7 | 13.6 | 15.7 | 28.4 | 21.6 | 21.6 |
| Self-Debug (Simple) ( Chen et al., 2024b ) | 16.7 | 19.4 | 17.1 | 23.2 | 19.7 | 19.6 |
| Self-Debug (Expl) ( Chen et al., 2024b ) | 22.2 | 17.5 | 15.7 | 25.3 | 22.1 | 22.6 |
| SQL-specific refinement | ||||||
| Qwen2.5-Instruct | Qwen3-8B (thinking) | ||||
| Method | 7B | 14B | 32B | OFF | ON |
| Base | 28.2 | 36.4 | 41.8 | 35.5 | 48.2 |
| Reasoning and self-correction | |||||
| CoT ( Kojima et al., 2022 ) | 32.7 | 40.9 | 42.7 | 39.1 | 52.7 |
| Self-Debug (Simple) ( Chen et al., 2024b ) | 35.5 | 45.5 | 50.9 | 40.9 | 48.2 |
| Self-Debug (Expl) ( Chen et al., 2024b ) | 28.2 | 38.2 | 45.5 | 45.5 | 48.2 |
| Method | 1 error | 2 errors | 3+ errors | Overall |
| Base | 28.2 | 23.2 | 16.9 | 22.6 |
| Reasoning and self-correction | ||||
| CoT ( Kojima et al., 2022 ) | 32.7 | 25.9 | 14.8 | 24.4 |
| Self-Debug (Simple) ( Chen et al., 2024b ) | 35.5 | 24.6 | 21.8 | 26.0 |
| Self-Debug (Expl) ( Chen et al., 2024b ) | 28.2 | 26.9 | 19.7 | 25.3 |
| SQL-specific refinement | ||||
| Method | Attribute | Value | Condition | Table | Others | Overall | ||||||
| Base | 24.3 | 18.9 | 41.4 | 37.9 | 29.4 | 29.4 | 33.3 | 46.7 | 16.7 | 0 8.3 | 30.0 | 28.2 |
| Reasoning and self-correction | ||||||||||||
| CoT | 35.1 | 21.6 | 37.9 | 34.5 | 52.9 | 47.1 | 26.7 | 53.3 | 16.7 | 16.7 | 35.5 | 32.7 |
| Self-Debug (Simple) | 24.3 | 24.3 | 41.4 | 48.3 | 35.3 | 41.2 | 53.3 | 46.7 | 16.7 | 16.7 | 33.6 | 35.5 |
| Self-Debug (Expl) | 24.3 | 16.2 | 37.9 | 31.0 | 35.3 | 35.3 | 33.3 | 53.3 | 0 8.3 | 16.7 | 29.1 | 28.2 |
| SQL-specific refinement | ||||||||||||
| 1 error | 2 errors | 3+ errors | Overall | |
| Exact match | 52.7 | 0 8.8 | 0 3.5 | 16.2 |
| Exact match, category only | 55.5 | 14.8 | 0 3.5 | 20.0 |
| Method | Attribute | Value | Condition | Table | Others | Overall |
| Qwen2.5-7B | 24.3 | 41.4 | 29.4 | 33.3 | 16.7 | 30.0 |
| TEG (predicted types) | 37.8 | 55.2 | 52.9 | 53.3 | 25.0 | 45.5 |
| TEG (oracle types) | 37.8 | 51.7 | 58.8 | 66.7 | 25.0 | 47.3 |
| Qwen2.5-14B | 27.0 | 37.9 | 41.2 | 40.0 | 41.7 | 35.5 |
| TEG (predicted types) | 35.1 | 58.6 | 47.1 | 80.0 | 25.0 | 48.2 |
| TEG (oracle types) | 48.6 | 55.2 | 58.8 | 66.7 | 33.3 | 52.7 |