Organizations: Fujian Key Laboratory of Urban Intelligent Sensing and Computing, Xiamen University, Xiamen, China · School of Computer, National University of Defense Technology, China
Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.
Figures & tables
Figure 1: Experiments with Llama-3.1-8B-Instruct on HumanEval. Left: the quality of the generated workflow for solving specific test problems. Middle: the number of executable workflows discovered during topology exploration on the training set. Right: Total token consumption.
Figure 2: The overall framework of our proposed FlowMAS.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
50.1 ±1.4
73.7 ±1.4
18.9 ±1.5
67.1 ±1.3
60.5 ±1.6
68.8 ±2.1
56.52
CoT [ 27 ]
51.2 ±1.5
75.6 ±1.1
21.0 ±1.7
69.4 ±1.4
64.0 ±1.9
73.1 ±1.9
59.05
CoT-SC [ 28 ]
52.4 ±0.8
77.5 ±0.9
22.1 ±0.0
66.8 ±0.8
66.0 ±0.9
75.4 ±1.3
60.03
ReAct [ 29 ]
51.9 ±1.9
75.2 ±1.7
28.1 ±2.0
65.4 ±1.5
63.7 ±1.7
71.9 ±1.7
59.37
Reflexion [ 30 ]
50.5 ±1.7
79.1 ±1.1
22.0 ±2.1
65.7 ±1.4
57.6 ±1.5
64.6 ±2.0
56.58
LLM-Majority [ 9 ]
56.6 ±0.0
76.7 ±1.2
22.4 ±1.6
67.3 ±0.9
62.0 ±1.1
71.4 ±1.2
59.40
Table 1: Main results on Llama-3.1-8B-Instruct. We bold the best results and underline the second-best results. Each experiment is run five times to record the mean and standard deviation.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
GPT-4o-mini
IO
86.3 ±0.6
86.3 ±0.5
45.0 ±0.7
71.2 ±0.4
66.7 ±0.6
73.2 ±0.9
71.45
ADAS [ 13 ]
83.4 ±1.2
84.9 ±0.9
42.7 ±1.2
67.3 ±0.8
63.9 ±0.9
75.5 ±1.2
69.62
AgentNet [ 34 ]
86.8 ±0.9
91.8 ±0.0
47.5 ±0.1
71.1 ±0.2
64.5 ±0.2
80.8 ±0.0
73.75
SELFORG [ 35 ]
86.0 ±2.0
92.7 ±0.1
50.7 ±0.7
71.7 ±0.3
68.9 ±0.0
84.3 ±0.1
75.72
AFlow [ 14 ]
89.6 ±0.7
90.5 ±0.7
50.5 ±0.9
81.0 ±0.6
72.6 ±0.9
79.7 ±0.5
77.32
Table 2: Results of IO and selected automated MAS-design methods on Gemma-4-26B-A4B-it and GPT-4o-mini backbones. We bold the best results and underline the second-best results.
Method
Acc. (%)
Training tokens
Train ×
Inference tokens
Infer ×
DyLAN [ 32 ]
51.9
5,842,507
15.2 ×
25,064,472
27.8 ×
ADAS [ 13 ]
59.5
84,526,998
220.0 ×
97,526,460
108.3 ×
AgentNet [ 34 ]
50.8
792,775
2.1 ×
451,047
0.5 ×
SELFORG [ 35 ]
61.9
853,528
2.2 ×
2,755,046
3.1 ×
AFlow [ 14 ]
65.1
6,253,861
16.3 ×
13,690,205
15.2 ×
MaAS [ 3 ]
62.7
223,042
0.6 ×
1,151,392
1.3 ×
Table 3: Efficiency comparison between FlowMAS and state-of-the-art baselines on the HumanEval. Train × and Infer × are relative to FlowMAS
Method
HumanEval
GSM8K
HotpotQA
ADAS [ 13 ]
59.5
79.5
60.9
FlowMAS
69.6
86.2
73.6
FlowMAS (VAE)
62.6
81.6
46.1
FlowMAS (Graph Diffusion Model)
55.7
80.8
50.6
FlowMAS (Transformer)
58.8
82.5
53.2
FlowMAS (MLP)
60.3
81.1
62.9
Table 4: Q1. Generator backbone ablation.
Variant
HE Acc.
G8K Acc.
HQ F1.
Avg. Topo. Count
ADAS [ 13 ]
59.5
79.5
60.9
31
w/o CDM
67.3
84.9
72.0
68
w/o CDM, w/ Graph Edit Distance [ 42 ]
66.8
85.2
72.1
73
w/o CDM, w/ Intrinsic Curiosity Module [ 43 ]
67.9
85.4
72.7
73
w/o IOM
65.0
84.2
69.3
60
w/o IOM, w/ Embedding Similarity
62.0
84.2
68.5
56
Table 5: Q2. Method-component ablation. HE, G8K, and HQ denote accuracy on HumanEval, GSM8K, and HotpotQA, respectively. “Avg. Topo. Count” denotes the average number of executable topologies discovered during training across the three benchmarks.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
rstep(t)←α⋅rcur(t)+β⋅∑Di(rpid(t)+rch(t))
Appendix
Algorithm 1 FlowMAS
Figure 3: Case study of the workflow topologies designed by AFlow, MaAS, and FlowMAS on DROP Benchmark.
Domain
Dataset
#Train/Val
#Test
Metric
Code Generation
HumanEval
33
131
pass@1
MBPP
86
341
pass@1
Math Reasoning
GSM8K
264
1,055
Solve rate
MATH
119
486
Solve rate
Reading Comprehension
HotpotQA
200
800
F1
DROP
200
800
F1
Appendix
Table 6: Dataset statistics used in our experiments.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
86.3 ±0.6
86.3 ±0.5
45.0 ±0.7
71.2 ±0.4
66.7 ±0.6
73.2 ±0.9
71.45
CoT [ 27 ]
86.6 ±0.7
86.4 ±0.5
45.8 ±0.6
72.6 ±0.6
66.4 ±0.9
77.0 ±0.8
72.47
CoT-SC [ 28 ]
87.6 ±0.0
86.5 ±0.3
46.5 ±0.0
74.2 ±0.5
67.4 ±0.6
77.5 ±0.7
73.28
ReAct [ 29 ]
87.0 ±0.8
86.5 ±0.7
43.3 ±0.9
77.7 ±0.6
71.0 ±0.8
73.8 ±0.9
73.22
Reflexion [ 30 ]
88.7 ±0.7
87.5 ±0.4
45.9 ±0.9
76.8 ±0.6
66.2 ±0.7
77.1 ±0.8
73.70
LLM-Majority [ 9 ]
87.9 ±0.3
86.5 ±0.6
46.8 ±0.0
74.4 ±0.0
68.3 ±0.6
74.6 ±0.5
73.08
Appendix
Table 7: Detailed results on GPT-4o-mini. The best result in each column is boldfaced, and the second-best is underlined. Each experiment is run five times to record the mean, standard deviation.
Method
HumanEval
GSM8K
MATH
MBPP
HotpotQA
DROP
Avg.
IO
90.8 ±0.8
90.5 ±0.7
61.3 ±0.9
73.3 ±0.6
74.9 ±0.8
83.9 ±1.0
79.12
CoT [ 27 ]
91.6 ±0.9
91.6 ±0.6
61.7 ±1.0
76.7 ±0.7
79.3 ±1.1
87.4 ±0.9
81.38
CoT-SC [ 28 ]
87.6 ±0.0
91.2 ±0.5
61.7 ±0.3
78.0 ±0.4
78.3 ±0.6
87.8 ±0.7
80.77
ReAct [ 29 ]
90.0 ±1.0
90.3 ±0.9
60.0 ±1.1
77.4 ±0.8
78.6 ±1.0
88.2 ±0.9
80.75
Reflexion [ 30 ]
91.8 ±0.9
92.1 ±0.6
62.2 ±1.2
76.2 ±0.7
73.5 ±0.8
85.1 ±1.0
80.15
LLM-Majority [ 9 ]
91.7 ±0.8
91.2 ±0.4
61.3 ±0.4
78.5 ±0.0
79.3 ±0.7
88.2 ±0.6
81.70
Appendix
Table 8: Detailed results on Gemma-4-26B-A4B-it. The best result in each column is boldfaced, and the second-best is underlined. Each experiment is run five times to record the mean, standard deviation.
# Operators
Method
Training Tokens
GPU Hours
Wall-clock Time (min)
Accuracy
10
AFlow
6,253,861
1.61
97
65.1
MaAS
223,042
0.52
31
62.7
FlowMAS
384,278
0.65
39
69.6
15
AFlow
9,746,315
2.90
174
61.8
MaAS
493,110
0.78
47
61.8
FlowMAS
427,057
0.70
42
68.7
Appendix
Table 9: Scalability comparison on HumanEval with different numbers of candidate operators.
Backbone
Method
Accuracy
Llama-3.1-8B-Instruct
AFlow
43.7
MaAS
60.0
FlowMAS
69.0
GPT-4o-mini
AFlow
75.7
MaAS
81.6
FlowMAS
85.0
Appendix
Table 10: Generalization results on BIG-Bench Hard across different LLM backbones.
α
0.1
0.2
0.3
0.4
0.5
Accuracy
67.9
69.6
69.6
68.7
67.9
β
0.6
0.7
0.8
0.9
1.0
Accuracy
66.4
68.7
69.6
67.9
67.9
Appendix
Table 11: Hyperparameter sensitivity analysis on HumanEval with Llama-3.1-8B-Instruct.
Method
Training Time (min)
Inference Time (min)
Accuracy
ADAS
200
91
59.5
AFlow
97
66
65.1
MaAS
31
22
62.7
FlowMAS
39
13
69.6
Appendix
Table 12: End-to-end training and inference wall-clock time on HumanEval with Llama-3.1-8B-Instruct. All experiments are measured on a single NVIDIA GeForce RTX 5090.
Method
Total Training
IOM Time
CDM Time
Embedding Time
Time (min)
(min)
(min)
(min)
FlowMAS
39
1.81
0.74
2.89
Appendix
Table 13: Detailed training-time overhead of FlowMAS on HumanEval. IOM Time includes the embedding computation required for the IOM reward, while Embedding Time denotes the total embedding-model computation throughout training.
Large language model (LLM)-based multi-agent systems have shown strong potential on complex tasks through agent specialization, tool use, and collaborative reasoning. However, most automated multi-agent system design methods still follow a one-shot paradigm: a workflow is optimized or selected before execution and then reused unchanged throughout the task. This static coordination strategy is ill-suited for long-horizon tasks whose subgoals, intermediate evidence, and information needs evolve over multiple execution stages. We propose EvoMAS, a framework for execution-time multi-agent workflow construction. EvoMAS formulates workflow construction as a meta-level sequential decision problem along a single task trajectory. At each stage, it constructs an explicit task state through a Planner-Evaluator-Updater pipeline and uses a learned Workflow Adapter to instantiate a stage-specific layered workflow from a fixed pool of candidate agents. The adapter is trained with policy gradients using sparse, verifiable terminal task success as the main supervision signal, while evaluator-based process reward is analyzed separately under very-hard sparse-reward settings. Experiments on GAIA, HLE, and DeepResearcher show that EvoMAS outperforms single-agent baselines and recent automated multi-agent workflow design methods. Our analyses further show that explicit task-state construction and learned workflow adaptation provide complementary benefits. Additional results indicate that process reward is most useful when terminal success is extremely sparse, and qualitative case studies illustrate that EvoMAS adapts agent coordination as the task state evolves.
Chengdong Xu, Kaiqiang Ke, Ziheng Liu +4
Sun Yat-Sen University · ShanghaiTech University · Zhejiang University +1
Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where performance critically depends on the underlying inter-agent communication topology. However, existing topology generation methods mainly optimize for isolated tasks, while real-world deployments involve streams of evolving tasks, requiring previously effective collaboration patterns to be retained and reused rather than rediscovered or overwritten. We identify a previously underexplored failure mode, \emph{topology forgetting}, in which adapting to new tasks shifts the topology generator away from communication structures required by earlier tasks. This issue stems from cross-task misalignment in both agent-level functional semantics and relational communication structures. To address this challenge, we propose \textbf{\textsc{MasFACT}}, a geometry-aware posterior transfer framework that preserves and reuses historical collaboration knowledge as transferable topology priors. We transfer these priors across task-specific agent spaces through Fused Gromov-Wasserstein optimal transport and perform PAC-Bayes-guided conservative posterior adaptation to balance task-specific plasticity with structural stability. Experiments across class-, domain-, and task-level continual settings demonstrate that \textsc{MasFACT} consistently improves average accuracy while reducing topology forgetting compared to strong topology generation and replay-based baselines, and can be seamlessly integrated with different MAS topology generators.
Xuefei Wang, Jialu Wang, Fengbo Zhang +6
Beihang University · Independent Researcher · Beijing University of Posts and Telecommunications
Multi-agent systems provide a powerful way to extend large language models (LLMs) by decomposing a complex task into specialized subtasks handled by different agents. However, their performance is often hindered by error propagation, arising from suboptimal workflow design or inaccurate agent outputs, which can propagate through the agent collaboration process and degrade final results. To address the challenges, we present MANGO (Multi-Agent Network Gradient Optimization), a data-driven framework that organizes and refines agent collaboration via a flow network constructed from past successful workflows. MANGO integrates reinforcement learning and textual gradients to jointly optimize workflow paths and agent behaviors, while a skipping mechanism prevents redundant updates to well-optimized agents for improving efficiency. Extensive experiments on seven benchmarks show that MANGO achieves up to 12.8% performance improvement over state-of-the-art baselines, enhances efficiency by 47.4%, and generalizes effectively to unseen domains. Our code and datasets are publicly available at https://github.com/openJiuwen-ai/agent-store/tree/main/community/mango.