CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence
Organizations: Peking University · National Key Lab of Data Space Technology and System · Baidu Inc.
Abstract
Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at https://github.com/BarryQ/CypherTurn.
Figures & tables
| Benchmark | MT | Cypher | Ctx | Agent | T/S |
|---|---|---|---|---|---|
| CypherBench | ✗ | ✓ | ✗ | ✗ | 1.0 |
| ZOGRASCOPE | ✗ | ✓ | ✗ | ✗ | 1.0 |
| SM3-Text-to-Query | ✗ | ✓ | ✗ | ✗ | 1.0 |
| Text2GQL-Bench | ✗ | ✓ | ✗ | ✗ | 1.0 |
| MTGQL † | ✓ | ✗ | ✗ | ✗ | 6.5 |
| CypherTurn | ✓ | ✓ | ✓ | ✓ | 8.2 |
| Guided Protocol | Agentic 3 | ||||||||||
| Model | EX | PSJS | CER | SEM | Tok/S | EX | PSJS | CER | SEM | Tok/S | EX |
| Frontier Models | |||||||||||
| Claude Opus 4.7 | 0.647 | 0.846 | 0.434 | 0.046 | 617 | 0.402 | 0.681 | 0.712 | 0.019 | 639 | 0.245 |
| GPT-5.5 | 0.625 | 0.828 | 0.424 | 0.025 | 504 | 0.401 | 0.673 | 0.698 | 0.006 | 554 | 0.224 |
| Kimi-K2.5 | 0.602 | 0.819 | 0.469 | 0.028 | 753 | 0.335 | 0.612 | 0.749 | 0.006 | 965 | 0.267 |
| Gemini-3.1-Flash-Lite | 0.572 | 0.744 | 0.508 | 0.017 | 500 | 0.432 | 0.698 | 0.695 | 0.008 | 427 | 0.140 |
| Model | INSPECT% | Agentic EX | EX |
|---|---|---|---|
| GPT-5.5 | 94.1 | 0.401 | 0.224 |
| Claude Opus 4.7 | 85.8 | 0.402 | 0.245 |
| Gemini-3.1-FL | 83.5 | 0.432 | 0.140 |
| GLM-5 | 54.3 | 0.274 | 0.232 |
| STRuCT-LLM-Novo | 40.4 | 0.333 | 0.193 |
| Kimi-K2.5 | 34.8 | 0.335 | 0.267 |
| Model | Paradigm | Guided | Agentic |
|---|---|---|---|
| Gemma-2-9B | Base | 0.281 | 0.136 |
| text-to-cypher-gemma | Single-turn SFT | 0.261 | 0.143 |
| QwQ-32B ‡ | Base | 0.395 | 0.324 |
| STRuCT-LLM-Novo | RL + CoT | 0.526 | 0.333 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Phenomenon | Category | Chain Mode | Dep. | Definition and Cypher Pattern |
|---|---|---|---|---|
| EXPAND | Navigation | expand | ✓ | Traverse one hop outward from the prior result set; the prior result acts as the seed node set. MATCH (n)-[:R]->(m) WHERE n.name IN $prev |
| PIVOT | Navigation | pivot | ✓ | Rank the prior result by a property, then traverse from the top- entities. WITH n ORDER BY n.prop DESC LIMIT k MATCH (n)-[:R]->(m) |
| FIRST | Navigation | pivot | ✓ | Reference the single top-ranked entity from the prior result and traverse from it. WITH n ORDER BY n.prop DESC LIMIT 1 MATCH (n)-[:R]->(m) |
| TOPIC_SHIFT | Discourse | topic_shift | ✗ | A context-break: the user poses an entirely new, independent query unrelated to prior results. The only non-chain-dependent phenomenon. |
| AGG_AVG | Aggregation | agg_numeric | ✓ | Compute the mean of a numeric property over the prior result set. MATCH (n) WHERE n.name IN $prev RETURN avg(n.prop) |
| AGG_MAX | Aggregation | agg_numeric | ✓ | Find the maximum value of a numeric property in the prior result set. MATCH (n) WHERE n.name IN $prev RETURN max(n.prop) |
| Chain Mode | Dep. | Definition and Example |
|---|---|---|
| None (T1) | ✗ | First turn of a session; no prior result exists. The query is fully self-contained. E.g. “List all provinces.” MATCH (p:Province) RETURN p.name |
| topic_shift | ✗ | A deliberate context-break mid-session. The user abandons the prior thread and opens a new independent query. E.g. (after querying officials) “Actually, list all trade routes.” |
| expand | ✓ | Traverse one hop outward from the prior result. The prior entity set is used directly as a seed. E.g. (after T1 provinces) “What cities are in those provinces?” |
| pivot | ✓ | Rank the prior result by a property, select top- , and traverse from those entities. E.g. “Among those cities, which one has the highest population? Show its legions.” |
| agg_numeric | ✓ | Compute a numeric aggregate (avg, max, sum) over a property of the prior result set. E.g. “What is the average tribute gold of those provinces?” |
| aggregate | ✓ | Count the cardinality of the prior result set. E.g. “How many legions were returned above?” |
| Persona | Definition | Example Utterance |
|---|---|---|
| Journalist | Investigative, direct phrasing; seeks evidence and facts; often frames queries as questions with implicit causal intent. | “Which provinces experienced the sharpest decline in tribute payments last year?” |
| Student | Exploratory and learning-oriented; may ask elementary questions or seek clarification on results. | “I’m trying to understand the hierarchy — can you show me which officials answer to the Emperor?” |
| Product Manager | Goal-focused; wants summary statistics and ranked lists; minimal tolerance for schema details. | “Give me the top three trade routes by cargo volume.” |
| Novice User | Vague phrasing; often uses pronouns ambiguously (“those things”, “that one”) without clear referents. | “Now what about the other ones? The ones with the big numbers?” |
| Data Analyst | Precise and technical; requests specific columns, sorted output, and numeric breakdowns. | “Return the entity ID, name, and tribute_gold for each province, ordered descending.” |
| Casual User | Informal register; contractions, colloquialisms, and incomplete sentences. May mix social and query intent. | “ok cool — so like, how many of those are there actually?” |
| Criterion | |
|---|---|
| Gold Cypher captures intent | 0.91 |
| Cross-turn anaphora unambiguous | 0.86 |
| Utterance natural for assigned persona | 0.64 |
| Pooled (reported) | 0.83 |
| Action | Input | Returns |
|---|---|---|
| EXECUTE_CYPHER | Cypher query string | Result rows; error on failure |
| INSPECT_SCHEMA | Empty (ignored) | Node labels, properties, relation types |
| SEARCH_VALUES | JSON: label, property, query | Up to 10 matching values |
| ASK_USER | Natural language question | String reply from user |
| SUBMIT_ANSWER | Final Cypher query | "Answer submitted" |
| Model | Anc | Mag | Ocn | Stl | Arc | Cel | Mer | Avg |
|---|---|---|---|---|---|---|---|---|
| Frontier Models | ||||||||
| Claude Opus 4.7 | 0.674 | 0.594 | 0.660 | 0.603 | 0.683 | 0.663 | 0.650 | 0.647 |
| GPT-5.5 | 0.633 | 0.570 | 0.659 | 0.597 | 0.671 | 0.667 | 0.581 | 0.625 |
| Kimi-K2.5 | 0.620 | 0.568 | 0.588 | 0.551 | 0.667 | 0.618 | 0.605 | 0.602 |
| Gemini-3.1-Flash-Lite | 0.574 | 0.620 | 0.535 | 0.520 | 0.604 | 0.595 | 0.555 | 0.572 |
| Qwen3-235B | 0.548 | 0.488 | 0.529 | 0.472 | 0.535 | 0.564 | 0.457 | 0.513 |
| Model | Anc | Mag | Ocn | Stl | Arc | Cel | Mer | Avg |
|---|---|---|---|---|---|---|---|---|
| Frontier Models | ||||||||
| Claude Opus 4.7 | 0.427 | 0.376 | 0.404 | 0.358 | 0.410 | 0.454 | 0.384 | 0.402 |
| GPT-5.5 | 0.433 | 0.323 | 0.494 | 0.389 | 0.416 | 0.415 | 0.340 | 0.401 |
| Kimi-K2.5 | 0.363 | 0.275 | 0.339 | 0.278 | 0.373 | 0.411 | 0.311 | 0.335 |
| Gemini-3.1-Flash-Lite | 0.445 | 0.479 | 0.442 | 0.392 | 0.409 | 0.446 | 0.409 | 0.432 |
| Qwen3-235B | 0.221 | 0.167 | 0.206 | 0.166 | 0.184 | 0.248 | 0.204 | 0.199 |
| Phenom. | Cat. | Cl | GPT | Km | Gfl | ST | Q3 | GL | DS | ER | Mi | Mx | Ll | G2 | t2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AGG_MAX | Agg | 0.842 | 0.813 | 0.788 | 0.741 | 0.796 | 0.798 | 0.840 | 0.793 | 0.603 | 0.810 | 0.310 | 0.360 | 0.409 | 0.197 | 0.650 |
| AGG_AVG | Agg | 0.767 | 0.679 | 0.740 | 0.677 | 0.740 | 0.722 | 0.780 | 0.700 | 0.578 | 0.702 | 0.285 | 0.336 | 0.422 | 0.298 | 0.602 |
| TOPIC_SHIFT | Disc | 0.821 | 0.849 | 0.760 | 0.782 | 0.680 | 0.643 | 0.487 | 0.656 | 0.778 | 0.422 | 0.111 | 0.788 | 0.538 | 0.673 | 0.642 |
| AGG_SUM | Agg | 0.642 | 0.663 | 0.661 | 0.655 | 0.665 | 0.692 | 0.696 | 0.676 | 0.547 | 0.661 | 0.281 | 0.331 | 0.397 | 0.318 | 0.563 |
| VALUE_FILTER | Filt | 0.725 | 0.753 | 0.764 | 0.590 | 0.657 | 0.691 | 0.562 | 0.702 | 0.590 | 0.528 | 0.112 | 0.270 | 0.354 | 0.382 | 0.549 |
| COUNT | Agg | 0.670 | 0.614 | 0.636 | 0.583 | 0.514 | 0.657 | 0.673 | 0.604 | 0.520 | 0.611 | 0.274 | 0.156 | 0.352 | 0.240 | 0.507 |
| Phenom. | Cat. | Cl | GPT | Km | Gfl | ST | Q3 | GL | DS | ER | Mi | Mx | Ll | G2 | t2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AGG_MAX | Agg | 0.663 | 0.611 | 0.599 | 0.579 | 0.549 | 0.505 | 0.579 | 0.510 | 0.443 | 0.601 | 0.480 | 0.195 | 0.187 | 0.121 | 0.473 |
| AGG_AVG | Agg | 0.576 | 0.500 | 0.507 | 0.455 | 0.471 | 0.374 | 0.516 | 0.381 | 0.316 | 0.527 | 0.399 | 0.141 | 0.155 | 0.130 | 0.389 |
| TOPIC_SHIFT | Disc | 0.533 | 0.644 | 0.433 | 0.739 | 0.446 | 0.225 | 0.289 | 0.233 | 0.563 | 0.249 | 0.222 | 0.178 | 0.200 | 0.313 | 0.376 |
| AGG_SUM | Agg | 0.524 | 0.501 | 0.457 | 0.418 | 0.414 | 0.345 | 0.426 | 0.343 | 0.310 | 0.443 | 0.331 | 0.121 | 0.119 | 0.094 | 0.346 |
| VALUE_FILTER | Filt | 0.399 | 0.416 | 0.225 | 0.337 | 0.287 | 0.101 | 0.169 | 0.112 | 0.292 | 0.112 | 0.039 | 0.090 | 0.135 | 0.197 | 0.208 |
| COUNT | Agg | 0.548 | 0.455 | 0.308 | 0.246 | 0.333 | 0.109 | 0.402 | 0.100 | 0.062 | 0.324 | 0.146 | 0.000 | 0.003 | 0.000 | 0.217 |
| Chain Mode | Dep. | Cl | GPT | Km | Gfl | ST | Q3 | GL | DS | ER | Mi | Mx | Ll | G2 | t2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| None (T1) | ✗ | 0.519 | 0.447 | 0.533 | 0.619 | 0.338 | 0.207 | 0.229 | 0.212 | 0.498 | 0.132 | 0.049 | 0.523 | 0.368 | 0.351 | 0.359 |
| topic_shift | ✗ | 0.821 | 0.849 | 0.760 | 0.782 | 0.680 | 0.643 | 0.487 | 0.656 | 0.778 | 0.422 | 0.111 | 0.788 | 0.538 | 0.673 | 0.642 |
| agg_numeric | ✓ | 0.753 | 0.716 | 0.730 | 0.690 | 0.733 | 0.737 | 0.771 | 0.723 | 0.576 | 0.724 | 0.292 | 0.342 | 0.409 | 0.271 | 0.603 |
| aggregate | ✓ | 0.670 | 0.614 | 0.636 | 0.583 | 0.514 | 0.657 | 0.673 | 0.604 | 0.520 | 0.611 | 0.274 | 0.156 | 0.352 | 0.240 | 0.507 |
| value_narrow | ✓ | 0.650 | 0.655 | 0.652 | 0.473 | 0.572 | 0.500 | 0.441 | 0.500 | 0.433 | 0.382 | 0.072 | 0.193 | 0.332 | 0.305 | 0.440 |
| pivot | ✓ | 0.611 | 0.554 | 0.489 | 0.513 | 0.401 | 0.501 | 0.456 | 0.494 | 0.335 | 0.416 | 0.068 | 0.083 | 0.094 | 0.103 | 0.366 |
| Chain Mode | Dep. | Cl | GPT | Km | Gfl | ST | Q3 | GL | DS | ER | Mi | Mx | Ll | G2 | t2 | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| None (T1) | ✗ | 0.383 | 0.409 | 0.304 | 0.555 | 0.347 | 0.202 | 0.211 | 0.214 | 0.510 | 0.164 | 0.136 | 0.455 | 0.411 | 0.411 | 0.336 |
| topic_shift | ✗ | 0.533 | 0.644 | 0.433 | 0.739 | 0.446 | 0.225 | 0.289 | 0.233 | 0.563 | 0.249 | 0.222 | 0.178 | 0.200 | 0.313 | 0.376 |
| agg_numeric | ✓ | 0.584 | 0.534 | 0.517 | 0.479 | 0.474 | 0.404 | 0.503 | 0.407 | 0.353 | 0.519 | 0.399 | 0.150 | 0.152 | 0.114 | 0.399 |
| aggregate | ✓ | 0.548 | 0.455 | 0.308 | 0.246 | 0.333 | 0.109 | 0.402 | 0.100 | 0.062 | 0.324 | 0.146 | 0.000 | 0.003 | 0.000 | 0.217 |
| value_narrow | ✓ | 0.326 | 0.307 | 0.198 | 0.262 | 0.233 | 0.075 | 0.115 | 0.086 | 0.251 | 0.075 | 0.029 | 0.083 | 0.118 | 0.158 | 0.165 |
| pivot | ✓ | 0.321 | 0.246 | 0.294 | 0.424 | 0.259 | 0.147 | 0.227 | 0.149 | 0.209 | 0.203 | 0.037 | 0.055 | 0.055 | 0.053 | 0.191 |
| Model | Imp | Jou | Cas | Stu | PMg | NNS | DA | Nov | Adv | Vrb | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Frontier Models | |||||||||||
| Claude Opus 4.7 | 0.559 | 0.627 | 0.678 | 0.661 | 0.621 | 0.635 | 0.619 | 0.646 | 0.708 | 0.720 | 0.647 |
| GPT-5.5 | 0.569 | 0.624 | 0.652 | 0.646 | 0.622 | 0.629 | 0.603 | 0.602 | 0.611 | 0.698 | 0.625 |
| Kimi-K2.5 | 0.595 | 0.588 | 0.624 | 0.609 | 0.585 | 0.621 | 0.593 | 0.573 | 0.569 | 0.678 | 0.602 |
| Gemini-3.1-Flash-Lite | 0.516 | 0.574 | 0.591 | 0.552 | 0.553 | 0.567 | 0.555 | 0.561 | 0.583 | 0.676 | 0.572 |
| Qwen3-235B | 0.507 | 0.505 | 0.526 | 0.465 | 0.509 | 0.498 | 0.543 | 0.468 | 0.542 | 0.578 | 0.513 |
| Model | Imp | Jou | Cas | Stu | PMg | NNS | DA | Nov | Adv | Vrb | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Frontier Models | |||||||||||
| Claude Opus 4.7 | 0.298 | 0.411 | 0.423 | 0.383 | 0.397 | 0.319 | 0.384 | 0.417 | 0.470 | 0.519 | 0.402 |
| GPT-5.5 | 0.347 | 0.409 | 0.441 | 0.394 | 0.397 | 0.363 | 0.382 | 0.396 | 0.447 | 0.437 | 0.401 |
| Kimi-K2.5 | 0.295 | 0.331 | 0.378 | 0.324 | 0.319 | 0.328 | 0.379 | 0.273 | 0.363 | 0.369 | 0.335 |
| Gemini-3.1-Flash-Lite | 0.383 | 0.423 | 0.444 | 0.427 | 0.399 | 0.404 | 0.437 | 0.481 | 0.424 | 0.495 | 0.432 |
| Qwen3-235B | 0.145 | 0.207 | 0.214 | 0.160 | 0.195 | 0.147 | 0.268 | 0.193 | 0.208 | 0.254 | 0.199 |
| Agentic 3 | Agentic 5 | Agentic 10 | |||||||||||||
| Model | EX | PSJS | CER | SEM | Tok/S | EX | PSJS | CER | SEM | Tok/S | EX | PSJS | CER | SEM | Tok/S |
| Frontier Models | |||||||||||||||
| Claude Opus 4.7 | 0.402 | 0.681 | 0.712 | 0.019 | 639 | 0.429 | 0.714 | 0.697 | 0.025 | 620 | 0.424 | 0.706 | 0.704 | 0.019 | 634 |
| GPT-5.5 | 0.401 | 0.673 | 0.698 | 0.006 | 554 | 0.395 | 0.661 | 0.707 | 0.004 | 564 | 0.406 | 0.680 | 0.704 | 0.007 | 546 |
| Kimi-K2.5 | 0.335 | 0.612 | 0.749 | 0.006 | 965 | 0.346 | 0.621 | 0.734 | 0.008 | 1,006 | 0.355 | 0.628 | 0.731 | 0.012 | 934 |
| Gemini-3.1-Flash-Lite | 0.432 | 0.698 | 0.695 | 0.008 | 427 | 0.431 | 0.697 | 0.694 | 0.007 | 427 | 0.430 | 0.695 | 0.694 | 0.007 | 404 |
| Model | EXEC | INSPECT | SEARCH | ASK |
|---|---|---|---|---|
| Frontier Models | ||||
| Claude Opus 4.7 | 0.048 | 0.858 | 0.001 | 0.093 |
| GPT-5.5 | 0.011 | 0.941 | 0.002 | 0.046 |
| Gemini-3.1-Flash-Lite | 0.107 | 0.835 | 0.002 | 0.056 |
| Kimi-K2.5 | 0.631 | 0.348 | 0.002 | 0.019 |
| GLM-5 | 0.407 | 0.543 | 0.001 | 0.048 |
| Variant | Models | Top-1 | Gemini keeps top-1 | |
|---|---|---|---|---|
| Agentic-PS | 0.915 | 10 | GPT-5.5 (0.463) | No |
| Disclosed-Horizon | 0.879 | 10 | Claude (0.456) | No |
| Budget-Encouraging | 0.786 | 7 | GPT-5.5 (0.399) | No |
| Model | Base | Agentic-PS | |
|---|---|---|---|
| GPT-5.5 | 0.402 | 0.463 | 0.061 |
| Claude Opus 4.7 | 0.400 | 0.461 | 0.061 |
| Kimi-K2.5 | 0.327 | 0.426 | 0.099 |
| Gemini-3.1-Flash-Lite | 0.433 | 0.435 | 0.002 |
| GLM-5 | 0.281 | 0.316 | 0.035 |
| DeepSeek-V3.2 | 0.207 | 0.237 | 0.030 |
| Model | Base | Disclosed-Horizon | |
|---|---|---|---|
| Claude Opus 4.7 | 0.400 | 0.456 | 0.056 |
| GPT-5.5 | 0.402 | 0.395 | 0.008 |
| Gemini-3.1-Flash-Lite | 0.433 | 0.364 | 0.069 |
| Kimi-K2.5 | 0.327 | 0.317 | 0.010 |
| ERNIE-5.0 | 0.297 | 0.273 | 0.024 |
| GLM-5 | 0.281 | 0.252 | 0.028 |
| Model | Base | BEP 10 | |
|---|---|---|---|
| GPT-5.5 | 0.402 | 0.399 | 0.003 |
| Claude Opus 4.7 | 0.400 | 0.367 | 0.033 |
| Gemini-3.1-Flash-Lite | 0.433 | 0.354 | 0.079 |
| Kimi-K2.5 | 0.327 | 0.291 | 0.036 |
| GLM-5 | 0.281 | 0.200 | 0.081 |
| DeepSeek-V3.2 | 0.207 | 0.198 | 0.010 |
| Model | Base gap | Agentic-PS | DH | BEP |
|---|---|---|---|---|
| Gemini-3.1-Flash-Lite | 0.140 | 0.137 | 0.208 | 0.218 |
| Claude Opus 4.7 | 0.245 | 0.186 | 0.191 | 0.280 |
| Qwen3-235B | 0.314 | 0.292 | 0.316 | 0.285 |
| DeepSeek-V3.2 | 0.302 | 0.269 | 0.260 | 0.308 |
| Frontier spread | 0.174 | 0.155 | 0.125 | 0.090 |
| Model | Exhaust. (%) | EX (orig.) | EX (excl. cliff) | |
|---|---|---|---|---|
| Claude Opus 4.7 | 0.0 | 0.402 | 0.402 | 0.000 |
| GPT-5.5 | 0.0 | 0.401 | 0.401 | 0.000 |
| Gemini-3.1-Flash-Lite | 0.3 | 0.432 | 0.432 | 0.000 |
| Kimi-K2.5 | 26.1 | 0.335 | 0.358 | 0.023 |
| Qwen3-235B | 62.6 | 0.199 | 0.231 | 0.032 |
| DeepSeek-V3.2 | 62.7 | 0.204 | 0.236 | 0.033 |