Flowchart-oriented dialogue (FOD) systems aim to guide users through multi-turn decision-making or operational procedures by following a domain-specific flowchart to achieve a task goal. In this work, we formalize flowchart reasoning in FOD as grounding user input to flowchart nodes at each dialogue turn while ensuring node transition is consistent with the correct flowchart path. Despite recent advances of LLMs in task-oriented dialogue systems, adapting them to FOD still faces two limitations: (1) LLMs lack an explicit mechanism to represent and reason over flowchart topology, and (2) they are prone to hallucinations, leading to unfaithful flowchart reasoning. To address these limitations, we propose FloCA, a zero-shot flowchart-oriented conversational agent. FloCA uses an LLM for intent understanding and response generation, while delegating flowchart reasoning to an external tool that performs topology-constrained graph execution, ensuring faithful and logically consistent node transitions across dialogue turns. We further introduce an evaluation framework with an LLM-based user simulator and five new metrics covering reasoning accuracy and interaction efficiency. Extensive experiments on FLODIAL and PFDial datasets highlight the bottlenecks of existing LLM/VLM-based methods and demonstrate the superiority of FloCA. The code and dataset are publicly available at https://github.com/Jinzi-Zou/FloCA-flowchart-reasoning.
Figures & tables
Figure 1: Comparison of workflows for the existing methods and FloCA. (a) relies on the intrinsic reasoning ability of LLMs or VLMs, which suffers from hallucination, and (b) leverages a flowchart reasoning tool to preserve topological integrity and ensure faithful and logically consistent reasoning.
Figure 2: An overview of FloCA. FloCA consists of two core components: an instruction-following LLM and a faithful flowchart reasoning tool. The left figure shows an example dialogue in troubleshooting, where each agent output corresponds to the result of the flowchart reasoning for a specific flowchart node or an answer from FAQ database. The right figure depicts the entire multi-turn reasoning process, with colors representing the graph functions and LLM reasoning processes, showing the input and output at each step and how reasoning is carried out.
Method
Initialization Metrics
Task Success Metrics
Efficiency Metrics
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Root-init ↑
Middle-init ↑
Overall ↑
SA-FloNet Raghu et al. (2022)
93.71
41.79
67.75
1.65
0.20
0.00
23.24
Graph Serialization Methods
Llama-3.3-70B
100.00
0.00
50.00
0.20
0.00
86.12
69.54
Claude Opus 4.1
98.57
23.13
60.85
83.43
80.12
0.00
0.38
GPT-4o
98.28
7.46
52.87
78.26
73.49
0.00
0.85
Table 1: In-domain FOD results on FLODIAL. All values are multiplied by 100 for better readability. “Root-init” and “Middle-init” represent the subsets where the ground-truth initial node is the root node (350 samples) and a non-root node (134 samples) in a flowchart, respectively, and “Overall” is computed on all test samples. “Task Success Metrics” serve as the primary indicators of overall model performance. The bold font highlights the best results, and the underline highlights the second-best results.
Figure 3: Flowchart reasoning accuracy around domain knowledge QA on FLODIAL. “GS”, “RAG+GS” and “VLMs” denote baselines of graph serialization methods, RAG-enhanced graph serialization methods, and visual language models.
Figure 4: Distribution of path-coverage relations on PFDial in the in-domain setting. Each value denotes the percentage of test samples belonging to the corresponding category.
Method
Reasoning Metrics
Efficiency Metrics
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Fine-tuned on PFDial
Llama-3.1-8B
100.00
92.63
75.35
0.00
0.77
Qwen2.5-7B
100.00
95.18
75.63
0.00
0.80
FloCA
Llama-3.1-8B-Instruct
100.00
92.91
83.56
0.00
10.66
Table 2: In-domain FOD results on PFDial. All values are multiplied by 100 for better readability. The bold font highlights the best results, and the underline highlights the second-best results.
Group
Method
TNGA (%)
PCA (%)
1
w/o self-loop mechanism
84.67
83.02
2
w/o edge matching
91.30
90.26
3
w/o FAQ interruption recovery
90.68
89.64
–
FloCA
91.92
90.88
Table 3: Ablation study of FloCA on FLODIAL using GPT-5.
Method
Failure Rate (%)
SA-FloNet Raghu et al. (2022)
12.21
Graph Serialization Methods
Llama-3.3-70B
2.89
Claude Opus 4.1
2.86
GPT-4o
0.82
GPT-5
5.17
Qwen3
8.48
Table 4: Task failure rates attributable to incorrect initial node grounding on FLODIAL.
Method
Failure Rate (%)
FloCA
Llama-3.3-70B
16.14
Claude Opus 4.1
8.69
GPT-4o
8.28
GPT-5
8.07
Qwen3
8.28
DeepSeek-R1
9.52
Table 5: Task failure rates attributable to LLM edge matching errors on FLODIAL.
Data Source
SE ↑
CE ↑
MSTTR ↑
MTLD ↑
HDD ↑
FLODIAL
Annotated Dialogues
3.61
0.13
0.91
37.09
14.48
User Simulator
3.79
0.16
0.92
43.57
16.46
PFDial
Annotated Dialogues
2.35
0.03
0.97
6.45
5.46
User Simulator
2.62
0.04
0.98
11.33
7.00
Table 6: Language diversity comparison of user utterances in the in-domain setting. Metrics are computed per dialogue turn and averaged over all test samples.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Datasets
Domain
Train Dialogs
Val Dialogs
Test Dialogs
Train flowcharts
Val flowcharts
Test flowcharts
Root-init
Middle-init
FLODIAL
In-domain
1,798
456
Root-init
Middle-init
12
12
12
350
134
Out-of-domain
1,786
454
root node
middle node
8
2
2
408
90
PFDial
In-domain
12,705
–
Root-init
Middle-init
440
–
80
Appendix
Table 7: A statistic summary of the datasets.
Method
Reasoning Metrics
Efficiency Metrics
Failure ↓
INGA ↑
TNGA ↑
PCA ↑
TR ↓
NSR ↓
Fine-tuned on PFDial
Llama-3.1-8B
8.17
2.88
0.00
53.12
26.52
48.31
Qwen2.5-7B
99.27
66.34
37.50
13.22
3.66
0.72
FloCA
Llama-3.1-8B-Instruct
100.00
70.19
29.32
0.00
4.86
0.00
Appendix
Table 8: Out-of-domain FOD results on PFDial. The bold font highlights the best results, and the underline highlights the second-best results.
Error Type
Description
Proportion (%)
Examples
Ambiguous or indirect user response
The user utterance is ambiguous, incomplete, or does not directly answer the agent’s question.
40.0
Agent: “Can you check if a fusible link has blown?”; User: “Yes, I can check that. I’ve done it before on this vehicle.”; Predicted edge: “yes”; Ground truth: “no”
Partial semantic understanding
The user utterance contains sufficient information for correct edge matching, but the LLM fails to correctly interpret semantic cues such as negation, polarity, or logical consistency across multiple sentences.
48.0
Agent: “Can you check if a fusible link has blown?”; User: “There does not appear to be a blown fusible link. The fusible link appears okay.”; Predicted edge: “yes”; Ground truth: “no”
Negative-question polarity error
The LLM incorrectly maps the response polarity because it over-relies on surface negation words (e.g., “not”, “nothing”, and “do not”).
12.0
Agent: “When your Windows boots, is nothing being displayed on the LCD?”; User: “I do not see anything on it.”; Predicted edge: “no”; Ground truth: “yes”
Appendix
Table 9: Taxonomy of LLM edge matching errors made by FloCA with GPT-5 on FLODIAL. The proportions are computed over all identified edge-matching failures.
Data Source
SE ↑
CE ↑
MSTTR ↑
MTLD ↑
HDD ↑
PFDial
Annotated Dialogues
2.35
0.03
0.97
6.45
5.46
User Simulator
2.62
0.04
0.98
11.33
7.00
Appendix
Table 10: Language diversity comparison of user utterances in the out-of-domain setting. Metrics are computed per dialogue turn and averaged over all test samples.
Figure 5: Distribution of path-coverage relations on the PFDial in the out-of-domain setting.
Figure 6: The prompt template for user-known fact extraction.
Figure 7: The prompt template for user simulator.
Figure 8: The prompt template for interactive flowchart reasoning in FLODIAL.
Figure 9: The prompt of LLM judge for user simulator evaluation.
Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model's own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.
Faissal Izermine, Hanru Bai, Oscar Davis +1
Max Planck Institute for Intelligent Systems · ELLIS Institute Tübingen · ETH Zurich +3
In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malfunctions, we propose a novel textual representation of fault trees. Building on it, we construct a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments, evaluating a model's ability to assist in malfunction localization, which contains 3130 entries and 40.75 turns per entry on average. We train an end-to-end model to generate vague information to reflect user behavior and introduce long-range rollback and recovery procedures to simulate user error scenarios, enabling assessment of a model's integrated capabilities in task tracking and error recovery, and Gemini 2.5 pro archives the best performance.
Large language models (LLMs) have become increasingly used for various tasks, often coupled with Chain-of-Thought (CoT) prompting to boost accuracy. Recent work has shown that high label-prediction accuracy does not guarantee correct intermediate reasoning, and the causes of reasoning flaws vary from sample to sample, yet existing remedies either focus on a single domain or assume that one flaw type applies uniformly across samples. A simple mitigation method is to provide the model with the correct answer, but we show that this yields no consistent improvement in reasoning quality. This indicates that the problem cannot be fixed by LLMs' awareness of answers, and must instead be addressed through the structure of reasoning. Motivated by this, we propose CRAFT (Consensus Reasoning-knowledge-graph Aggregation for Flaw-aware Trace synthesis), which aggregates the consensus components shared across multiple candidate reasoning traces to synthesize improved ones. CRAFT consistently improves label-prediction accuracy on both logical and mathematical reasoning benchmarks, outperforming most baselines, while its post-processed traces achieve higher quality under fine-grained benchmark evaluation.
Zipeng Ling, Shuliang Liu, Seonil Son +4
The Hong Kong University of Science and Technology (Guangzhou) · University of Alberta · Alberta Machine Intelligence Institute +3