LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.
Figures & tables
Benchmark
Incomplete Specifications
Autonomous Requirement Recovery
Multi-Domain
AgentBench ( Liu et al., 2024 )
✘
✘
✔
GAIA ( Mialon et al., 2024 )
✘
✘
✔
MultiAgentBench ( Zhu et al., 2025 )
✘
✘
✔
SWE-Together ( Wu et al., 2026 )
✔
✘
✘
ICAE-Bench ( Peng et al., 2026 )
✔
✘
✘
SpecBench ( Hamblin et al., 2026 )
✔
✔
✘
Table 1: Comparison between ProgSpec and related agent benchmarks. Autonomous requirement recovery denotes recovery from reasoning or execution evidence rather than requirements supplied through user clarification.
Domain
Tasks
Reqs.
L1
L2
L3
Reqs./task
Coding
90
539
90
164
285
5.99
Math
90
530
180
213
137
5.89
QA
90
482
90
97
295
5.36
Total
270
1,551
360
474
717
5.74
Table 2: Statistics of ProgSpec. L1, L2, and L3 denote explicit, implicit, and latent requirements, respectively.
Method
HumanEval
HotpotQA
DROP
GSM8K
MBPP
ProgSpec
GPT-4o-mini
87.0
68.1
68.3
92.7
71.8
38.3
MetaGPT ( Hong et al., 2024 )
87.0
67.3
77.3
82.8
66.6
8.9
AgentCoder ( Huang et al., 2023 )
82.4
54.9
82.0
88.5
59.2
12.8
G-Designer ( Zhang et al., 2024 )
88.6
68.1
71.5
95.1
71.3
11.1
AFlow ( Zhang et al., 2025c )
94.7
73.5
80.6
93.5
83.4
39.9
Meta-Agent ( Xu and Tai, 2026 )
96.0 ∗
69.5 ∗
82.7 ∗
93.7 ∗
84.6 ∗
42.9
Table 3: Main results with GPT-4o-mini on six benchmarks. We report pass@1 for HumanEval and MBPP, F1 for HotpotQA and DROP, solve rate for GSM8K, and weighted requirement coverage for ProgSpec. Bold marks the best reported result in each column. * indicates results are reported in the original paper. The overall performance trend is statistically significant ( p=0.024 ).
Setting
ProgSpec
HumanEval
HotpotQA
RepoMAS
44.7
99.2
75.5
w/o Structured Issue
30.8
84.7
68.5
w/o Patch Validation
33.4
84.0
59.4
w/o Structural Revision
30.3
87.8
56.4
w/o Local Re-execution
32.4
84.7
69.7
Table 4: Ablation study of RepoMAS. Each variant removes one major component from the full framework.
Issue Representation
ProgSpec
HumanEval
HotpotQA
Avg.
Direct Revision
30.2
90.8
68.5
63.2
Free-form Issue
30.8
84.7
68.5
61.4
Structured Issue w/o Evidence
27.5
90.8
69.0
62.5
Full Structured Issue
44.7
99.2
75.5
73.1
Table 5: Effect of different Issue representations. All variants differ only in how execution feedback is represented before patch generation.
Metric
Value
Validation Precision
62.5%
Validation Recall
86.2%
Accepted-patch Error Rate
37.5%
Regression Rate
22.5%
Good patches: initial
13.8%
Good patches: revised
69.0%
Table 6: Effectiveness of Patch Validation on ProgSpec-Coding.
Variant
ProgSpec
HumanEval
HotpotQA
Full RepoMAS
44.7
99.2
75.5
w/o Task Graph Revision
25.8
90.1
55.8
w/o Agent Graph Revision
25.8
88.5
56.9
w/o Structural Revision
30.3
87.8
56.4
Table 7: Effects of Task and Agent Graph revision. Each variant disables one or both forms of structural revision.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Example and evidence
L1 (Explicit)
The user asks the function to return a JSON object. This requirement is stated directly in the request.
L2 (Implicit)
The task asks for the real-valued expression log(x−1) to be evaluated. The condition x>1 follows directly from the domain of the real logarithm, although it is not stated in the request.
L3 (Latent)
The request does not specify how empty input should be handled. Inspection of an available repository test shows that empty input must return an empty result rather than raise an exception.
Appendix
Table 8: Examples of ProgSpec requirement levels.
Backbone
Method
HumanEval
HotpotQA
DROP
GSM8K
MBPP
ProgSpec
GPT-5.5
GPT-5.5
99.24
59.38
39.12
93.46
83.28
92.08
MetaGPT ( Hong et al., 2024 )
98.47
32.88
3.38
71.18
81.23
89.95
AgentCoder ( Huang et al., 2023 )
99.24
50.62
33.75
92.61
83.58
93.02
G-Designer ( Zhang et al., 2024 )
100.00
57.00
36.50
93.36
84.46
94.20
AFlow ( Zhang et al., 2025c )
99.24
57.75
39.12
94.22
83.58
93.34
Meta-Agent ( Xu and Tai, 2026 )
98.47
54.37
25.50
92.89
80.06
89.50
Appendix
Table 9: Backbone robustness results on six benchmarks. We repeat the main experiments using different backbone models while keeping the methods and evaluation settings unchanged. All methods within each group use the same backbone model. Bold marks the highest score in each column.
Metric
Score
95% CI
Evaluation Unit
Annotation–review agreement
91.4% ( AC1=0.91 )
[88.0, 93.9]
329 / 360 judgments
Requirement validity
95.4%
[93.2, 96.9]
477 / 500 requirements
Requirement independence
98.2%
[96.6, 99.1]
491 / 500 requirements
Evidence sufficiency
97.4%
[95.6, 98.5]
487 / 500 requirements
Ambiguity rate
0.2%
[0.0, 1.1]
1 / 500 requirements
Appendix
Table 10: Human validation of ProgSpec annotations. Results are computed over 90 tasks and 500 annotated requirements.
Benchmark
Temperature
Max Tokens
Max Iterations
HumanEval
0.2
4096
3
HotpotQA / DROP / GSM8K / MBPP
0.2
4096
3
ProgSpec–Coding
0.0
4096
3
ProgSpec–Math / QA
0.0
4096
3
Appendix
Table 11: Evaluation settings across benchmarks. We report the temperature, maximum number of generated tokens, and maximum number of iterations used for each benchmark.
LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.
Yapeng Li, Songze Li, Shuang Yu +4
Harbin Institute of Technology, Harbin, China · Peking University, Beijing, China
MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Constructing the specification of a MAS requires formal language expertise, while machine learning techniques depend on labelled data which are rarely available. To address these issues, we propose genRTEC', a method that leverages pre-trained Large Language Models (LLMs) to generate executable MAS specifications, in the language of the Run-Time Event Calculus' (RTEC), from natural language descriptions. genRTEC constructs MAS specifications with complex hierarchical and cyclic dependencies based only on short natural language descriptions of the concepts involved. We present an extensive empirical evaluation of genRTEC, spanning various MAS specifications, including both a qualitative and a quantitative assessment. Our results demonstrate that genRTEC constructs executable MAS specifications of high predictive accuracy without compromising reasoning efficiency.
Andreas Kouvaras, Periklis Mantenoglou, Alexander Artikis
University of Piraeus, Greece. · Örebro University, Sweden. · NCSR “Demokritos”, Greece.
Large language model (LLM)-based Multi-agent systems (MAS) have shown promise in tackling complex collaborative tasks, where agents are typically orchestrated via role-specific prompts. While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. A core innovation of MASPO is its joint evaluation mechanism, which assesses prompts not merely by their local validity, but by their capacity to facilitate downstream success for successor agents. This effectively bridges the gap between local interactions and global outcomes without relying on ground-truth labels. Furthermore, MASPO employs a data-driven evolutionary beam search to efficiently navigate the high-dimensional prompt space. Extensive empirical evaluations across 6 diverse tasks demonstrate that MASPO consistently outperforms state-of-the-art prompt optimization methods, achieving an average accuracy improvement of 2.9. We release our code at https://github.com/wangzx1219/MASPO.
Zhexuan Wang, Xuebo Liu, Li Wang +4
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China.