Flowchart-oriented dialogue (FOD) systems aim to guide users through multi-turn decision-making or operational procedures by following a domain-specific flowchart to achieve a task goal. In this work, we formalize flowchart reasoning in FOD as grounding user input to flowchart nodes at each dialogue turn while ensuring node transition is consistent with the correct flowchart path. Despite recent advances of LLMs in task-oriented dialogue systems, adapting them to FOD still faces two limitations: (1) LLMs lack an explicit mechanism to represent and reason over flowchart topology, and (2) they are prone to hallucinations, leading to unfaithful flowchart reasoning. To address these limitations, we propose FloCA, a zero-shot flowchart-oriented conversational agent. FloCA uses an LLM for intent understanding and response generation, while delegating flowchart reasoning to an external tool that performs topology-constrained graph execution, ensuring faithful and logically consistent node transitions across dialogue turns. We further introduce an evaluation framework with an LLM-based user simulator and five new metrics covering reasoning accuracy and interaction efficiency. Extensive experiments on FLODIAL and PFDial datasets highlight the bottlenecks of existing LLM/VLM-based methods and demonstrate the superiority of FloCA. The code and dataset are publicly available at https://github.com/Jinzi-Zou/FloCA-flowchart-reasoning.
Figures & tables
Figure 1: Comparison of workflows for the existing methods and FloCA. (a) relies on the intrinsic reasoning ability of LLMs or VLMs, which suffers from hallucination, and (b) leverages a flowchart reasoning tool to preserve topological integrity and ensure faithful and logically consistent reasoning.
Figure 2: An overview of FloCA. FloCA consists of two core components: an instruction-following LLM and a faithful flowchart reasoning tool. The left figure shows an example dialogue in troubleshooting, where each agent output corresponds to the result of the flowchart reasoning for a specific flowchart node or an answer from FAQ database. The right figure depicts the entire multi-turn reasoning process, with colors representing the graph functions and LLM reasoning processes, showing the input and output at each step and how reasoning is carried out.
Method
Initialization Metrics
Task Success Metrics
Efficiency Metrics
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Root-init ↑
Middle-init ↑
Overall ↑
SA-FloNet Raghu et al. (2022)
93.71
41.79
67.75
1.65
0.20
0.00
23.24
Graph Serialization Methods
Llama-3.3-70B
100.00
0.00
50.00
0.20
0.00
86.12
69.54
Claude Opus 4.1
98.57
23.13
60.85
83.43
80.12
0.00
0.38
GPT-4o
98.28
7.46
52.87
78.26
73.49
0.00
0.85
Table 1: In-domain FOD results on FLODIAL. All values are multiplied by 100 for better readability. “Root-init” and “Middle-init” represent the subsets where the ground-truth initial node is the root node (350 samples) and a non-root node (134 samples) in a flowchart, respectively, and “Overall” is computed on all test samples. “Task Success Metrics” serve as the primary indicators of overall model performance. The bold font highlights the best results, and the underline highlights the second-best results.
Figure 3: Flowchart reasoning accuracy around domain knowledge QA on FLODIAL. “GS”, “RAG+GS” and “VLMs” denote baselines of graph serialization methods, RAG-enhanced graph serialization methods, and visual language models.
Figure 4: Distribution of path-coverage relations on PFDial in the in-domain setting. Each value denotes the percentage of test samples belonging to the corresponding category.
Method
Reasoning Metrics
Efficiency Metrics
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Fine-tuned on PFDial
Llama-3.1-8B
100.00
92.63
75.35
0.00
0.77
Qwen2.5-7B
100.00
95.18
75.63
0.00
0.80
FloCA
Llama-3.1-8B-Instruct
100.00
92.91
83.56
0.00
10.66
Table 2: In-domain FOD results on PFDial. All values are multiplied by 100 for better readability. The bold font highlights the best results, and the underline highlights the second-best results.
Group
Method
TNGA (%)
PCA (%)
1
w/o self-loop mechanism
84.67
83.02
2
w/o edge matching
91.30
90.26
3
w/o FAQ interruption recovery
90.68
89.64
–
FloCA
91.92
90.88
Table 3: Ablation study of FloCA on FLODIAL using GPT-5.
Method
Failure Rate (%)
SA-FloNet Raghu et al. (2022)
12.21
Graph Serialization Methods
Llama-3.3-70B
2.89
Claude Opus 4.1
2.86
GPT-4o
0.82
GPT-5
5.17
Qwen3
8.48
Table 4: Task failure rates attributable to incorrect initial node grounding on FLODIAL.
Method
Failure Rate (%)
FloCA
Llama-3.3-70B
16.14
Claude Opus 4.1
8.69
GPT-4o
8.28
GPT-5
8.07
Qwen3
8.28
DeepSeek-R1
9.52
Table 5: Task failure rates attributable to LLM edge matching errors on FLODIAL.
Data Source
SE ↑
CE ↑
MSTTR ↑
MTLD ↑
HDD ↑
FLODIAL
Annotated Dialogues
3.61
0.13
0.91
37.09
14.48
User Simulator
3.79
0.16
0.92
43.57
16.46
PFDial
Annotated Dialogues
2.35
0.03
0.97
6.45
5.46
User Simulator
2.62
0.04
0.98
11.33
7.00
Table 6: Language diversity comparison of user utterances in the in-domain setting. Metrics are computed per dialogue turn and averaged over all test samples.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Datasets
Domain
Train Dialogs
Val Dialogs
Test Dialogs
Train flowcharts
Val flowcharts
Test flowcharts
Root-init
Middle-init
FLODIAL
In-domain
1,798
456
Root-init
Middle-init
12
12
12
350
134
Out-of-domain
1,786
454
root node
middle node
8
2
2
408
90
PFDial
In-domain
12,705
–
Root-init
Middle-init
440
–
80
Appendix
Table 7: A statistic summary of the datasets.
Method
Reasoning Metrics
Efficiency Metrics
Failure ↓
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Fine-tuned on PFDial
Llama-3.1-8B
8.17
2.88
0.00
53.12
26.52
48.31
Qwen2.5-7B
99.27
66.34
37.50
13.22
3.66
0.72
FloCA
Llama-3.1-8B-Instruct
100.00
70.19
29.32
0.00
4.86
0.00
Appendix
Table 8: Out-of-domain FOD results on PFDial. The bold font highlights the best results, and the underline highlights the second-best results.
Error Type
Description
Proportion (%)
Examples
Ambiguous or indirect user response
The user utterance is ambiguous, incomplete, or does not directly answer the agent’s question.
40.0
Agent: “Can you check if a fusible link has blown?”; User: “Yes, I can check that. I’ve done it before on this vehicle.”; Predicted edge: “yes”; Ground truth: “no”
Partial semantic understanding
The user utterance contains sufficient information for correct edge matching, but the LLM fails to correctly interpret semantic cues such as negation, polarity, or logical consistency across multiple sentences.
48.0
Agent: “Can you check if a fusible link has blown?”; User: “There does not appear to be a blown fusible link. The fusible link appears okay.”; Predicted edge: “yes”; Ground truth: “no”
Negative-question polarity error
The LLM incorrectly maps the response polarity because it over-relies on surface negation words (e.g., “not”, “nothing”, and “do not”).
12.0
Agent: “When your Windows boots, is nothing being displayed on the LCD?”; User: “I do not see anything on it.”; Predicted edge: “no”; Ground truth: “yes”
Appendix
Table 9: Taxonomy of LLM edge matching errors made by FloCA with GPT-5 on FLODIAL. The proportions are computed over all identified edge-matching failures.
Data Source
SE ↑
CE ↑
MSTTR ↑
MTLD ↑
HDD ↑
PFDial
Annotated Dialogues
2.35
0.03
0.97
6.45
5.46
User Simulator
2.62
0.04
0.98
11.33
7.00
Appendix
Table 10: Language diversity comparison of user utterances in the out-of-domain setting. Metrics are computed per dialogue turn and averaged over all test samples.
Figure 5: Distribution of path-coverage relations on the PFDial in the out-of-domain setting.
Figure 6: The prompt template for user-known fact extraction.
Figure 7: The prompt template for user simulator.
Figure 8: The prompt template for interactive flowchart reasoning in FLODIAL.
Figure 9: The prompt of LLM judge for user simulator evaluation.