Organizations: Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University, Xiamen, China · School of Computer, National University of Defense Technology, China
Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.
Figures & tables
Figure 1: Experiments with Llama-3.1-8B-Instruct on HumanEval. Left: the quality of the generated workflow for solving specific test problems. Middle: the number of executable workflows discovered during topology exploration on the training set. Right: Total token consumption.
Figure 2: The overall framework of our proposed FlowMAS.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
50.1 ±1.4
73.7 ±1.4
18.9 ±1.5
67.1 ±1.3
60.5 ±1.6
68.8 ±2.1
56.52
CoT [ 27 ]
51.2 ±1.5
75.6 ±1.1
21.0 ±1.7
69.4 ±1.4
64.0 ±1.9
73.1 ±1.9
59.05
CoT-SC [ 28 ]
52.4 ±0.8
77.5 ±0.9
22.1 ±0.0
66.8 ±0.8
66.0 ±0.9
75.4 ±1.3
60.03
ReAct [ 29 ]
51.9 ±1.9
75.2 ±1.7
28.1 ±2.0
65.4 ±1.5
63.7 ±1.7
71.9 ±1.7
59.37
Reflexion [ 30 ]
50.5 ±1.7
79.1 ±1.1
22.0 ±2.1
65.7 ±1.4
57.6 ±1.5
64.6 ±2.0
56.58
LLM-Majority [ 9 ]
56.6 ±0.0
76.7 ±1.2
22.4 ±1.6
67.3 ±0.9
62.0 ±1.1
71.4 ±1.2
59.40
Table 1: Main results on Llama-3.1-8B-Instruct. We bold the best results and underline the second-best results. Each experiment is run five times to record the mean and standard deviation.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
GPT-4o-mini
IO
86.3 ±0.6
86.3 ±0.5
45.0 ±0.7
71.2 ±0.4
66.7 ±0.6
73.2 ±0.9
71.45
ADAS [ 13 ]
83.4 ±1.2
84.9 ±0.9
42.7 ±1.2
67.3 ±0.8
63.9 ±0.9
75.5 ±1.2
69.62
AgentNet [ 34 ]
86.8 ±0.9
91.8 ±0.0
47.5 ±0.1
71.1 ±0.2
64.5 ±0.2
80.8 ±0.0
73.75
SELFORG [ 35 ]
86.0 ±2.0
92.7 ±0.1
50.7 ±0.7
71.7 ±0.3
68.9 ±0.0
84.3 ±0.1
75.72
AFlow [ 14 ]
89.6 ±0.7
90.5 ±0.7
50.5 ±0.9
81.0 ±0.6
72.6 ±0.9
79.7 ±0.5
77.32
Table 2: Results of IO and selected automated MAS-design methods on Gemma-4-26B-A4B-it and GPT-4o-mini backbones. We bold the best results and underline the second-best results.
Method
Acc. (%)
Training tokens
Train ×
Inference tokens
Infer ×
DyLAN [ 32 ]
51.9
5,842,507
15.2 ×
25,064,472
27.8 ×
ADAS [ 13 ]
59.5
84,526,998
220.0 ×
97,526,460
108.3 ×
AgentNet [ 34 ]
50.8
792,775
2.1 ×
451,047
0.5 ×
SELFORG [ 35 ]
61.9
853,528
2.2 ×
2,755,046
3.1 ×
AFlow [ 14 ]
65.1
6,253,861
16.3 ×
13,690,205
15.2 ×
MaAS [ 3 ]
62.7
223,042
0.6 ×
1,151,392
1.3 ×
Table 3: Efficiency comparison between FlowMAS and state-of-the-art baselines on the HumanEval. Train × and Infer × are relative to FlowMAS
Method
HumanEval
GSM8K
HotpotQA
ADAS [ 13 ]
59.5
79.5
60.9
FlowMAS
69.6
86.2
73.6
FlowMAS (VAE)
62.6
81.6
46.1
FlowMAS (Graph Diffusion Model)
55.7
80.8
50.6
FlowMAS (Transformer)
58.8
82.5
53.2
FlowMAS (MLP)
60.3
81.1
62.9
Table 4: Q1. Generator backbone ablation.
Variant
HE Acc.
G8K Acc.
HQ F1.
Avg. Topo. Count
ADAS [ 13 ]
59.5
79.5
60.9
31
w/o CDM
67.3
84.9
72.0
68
w/o CDM, w/ Graph Edit Distance [ 42 ]
66.8
85.2
72.1
73
w/o CDM, w/ Intrinsic Curiosity Module [ 43 ]
67.9
85.4
72.7
73
w/o IOM
65.0
84.2
69.3
60
w/o IOM, w/ Embedding Similarity
62.0
84.2
68.5
56
Table 5: Q2. Method-component ablation. HE, G8K, and HQ denote accuracy on HumanEval, GSM8K, and HotpotQA, respectively. “Avg. Topo. Count” denotes the average number of executable topologies discovered during training across the three benchmarks.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
rstep(t)←α⋅rcur(t)+β⋅∑Di(rpid(t)+rch(t))
Appendix
Algorithm 1 FlowMAS
Figure 3: Case study of the workflow topologies designed by AFlow, MaAS, and FlowMAS on DROP Benchmark.
Domain
Dataset
#Train/Val
#Test
Metric
Code Generation
HumanEval
33
131
pass@1
MBPP
86
341
pass@1
Math Reasoning
GSM8K
264
1,055
Solve rate
MATH
119
486
Solve rate
Reading Comprehension
HotpotQA
200
800
F1
DROP
200
800
F1
Appendix
Table 6: Dataset statistics used in our experiments.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
86.3 ±0.6
86.3 ±0.5
45.0 ±0.7
71.2 ±0.4
66.7 ±0.6
73.2 ±0.9
71.45
CoT [ 27 ]
86.6 ±0.7
86.4 ±0.5
45.8 ±0.6
72.6 ±0.6
66.4 ±0.9
77.0 ±0.8
72.47
CoT-SC [ 28 ]
87.6 ±0.0
86.5 ±0.3
46.5 ±0.0
74.2 ±0.5
67.4 ±0.6
77.5 ±0.7
73.28
ReAct [ 29 ]
87.0 ±0.8
86.5 ±0.7
43.3 ±0.9
77.7 ±0.6
71.0 ±0.8
73.8 ±0.9
73.22
Reflexion [ 30 ]
88.7 ±0.7
87.5 ±0.4
45.9 ±0.9
76.8 ±0.6
66.2 ±0.7
77.1 ±0.8
73.70
LLM-Majority [ 9 ]
87.9 ±0.3
86.5 ±0.6
46.8 ±0.0
74.4 ±0.0
68.3 ±0.6
74.6 ±0.5
73.08
Appendix
Table 7: Detailed results on GPT-4o-mini. The best result in each column is boldfaced, and the second-best is underlined. Each experiment is run five times to record the mean, standard deviation.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
90.8 ±0.8
90.5 ±0.7
61.3 ±0.9
73.3 ±0.6
74.9 ±0.8
83.9 ±1.0
79.12
CoT [ 27 ]
91.6 ±0.9
91.6 ±0.6
61.7 ±1.0
76.7 ±0.7
79.3 ±1.1
87.4 ±0.9
81.38
CoT-SC [ 28 ]
87.6 ±0.0
91.2 ±0.5
61.7 ±0.3
78.0 ±0.4
78.3 ±0.6
87.8 ±0.7
80.77
ReAct [ 29 ]
90.0 ±1.0
90.3 ±0.9
60.0 ±1.1
77.4 ±0.8
78.6 ±1.0
88.2 ±0.9
80.75
Reflexion [ 30 ]
91.8 ±0.9
92.1 ±0.6
62.2 ±1.2
76.2 ±0.7
73.5 ±0.8
85.1 ±1.0
80.15
LLM-Majority [ 9 ]
91.7 ±0.8
91.2 ±0.4
61.3 ±0.4
78.5 ±0.0
79.3 ±0.7
88.2 ±0.6
81.70
Appendix
Table 8: Detailed results on Gemma-4-26B-A4B-it. The best result in each column is boldfaced, and the second-best is underlined. Each experiment is run five times to record the mean, standard deviation.
# Operators
Method
Training Tokens
GPU Hours
Wall-clock Time (min)
Accuracy
10
AFlow
6,253,861
1.61
97
65.1
MaAS
223,042
0.52
31
62.7
FlowMAS
384,278
0.65
39
69.6
15
AFlow
9,746,315
2.90
174
61.8
MaAS
493,110
0.78
47
61.8
FlowMAS
427,057
0.70
42
68.7
Appendix
Table 9: Scalability comparison on HumanEval with different numbers of candidate operators.
Backbone
Method
Accuracy
Llama-3.1-8B-Instruct
AFlow
43.7
MaAS
60.0
FlowMAS
69.0
GPT-4o-mini
AFlow
75.7
MaAS
81.6
FlowMAS
85.0
Appendix
Table 10: Generalization results on BIG-Bench Hard across different LLM backbones.
α
0.1
0.2
0.3
0.4
0.5
Accuracy
67.9
69.6
69.6
68.7
67.9
β
0.6
0.7
0.8
0.9
1.0
Accuracy
66.4
68.7
69.6
67.9
67.9
Appendix
Table 11: Hyperparameter sensitivity analysis on HumanEval with Llama-3.1-8B-Instruct.
Method
Training Time (min)
Inference Time (min)
Accuracy
ADAS
200
91
59.5
AFlow
97
66
65.1
MaAS
31
22
62.7
FlowMAS
39
13
69.6
Appendix
Table 12: End-to-end training and inference wall-clock time on HumanEval with Llama-3.1-8B-Instruct. All experiments are measured on a single NVIDIA GeForce RTX 5090.
Method
Total Training
IOM Time
CDM Time
Embedding Time
Time (min)
(min)
(min)
(min)
FlowMAS
39
1.81
0.74
2.89
Appendix
Table 13: Detailed training-time overhead of FlowMAS on HumanEval. IOM Time includes the embedding computation required for the IOM reward, while Embedding Time denotes the total embedding-model computation throughout training.