RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems
Organizations: Harbin Institute of Technology, Harbin, China
Abstract
LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool use, or execution. We refer to such problems as progressively specified tasks. To systematically study this setting, we introduce ProgSpec, a benchmark that evaluates final outputs against requirements explicitly stated in the initial request and additional requirements supported by the available task evidence. We further propose RepoMAS, an issue-driven multi-agent framework inspired by open-source project management. RepoMAS records newly discovered requirements, conflicts, and failures as structured Issues and uses them to revise the task specification and execution structure during problem solving. Across ProgSpec and five existing benchmarks, RepoMAS achieves the best performance. Further analyses show that its issue-driven revision and repository maintenance mechanisms consistently contribute to performance. These results highlight the importance of allowing MASs to revise not only how a task is solved, but also revise their explicit representation of task requirements during execution.
Figures & tables
| Benchmark | Incomplete Specifications | Autonomous Requirement Recovery | Multi-Domain |
|---|---|---|---|
| AgentBench ( Liu et al., 2024 ) | ✘ | ✘ | ✔ |
| GAIA ( Mialon et al., 2024 ) | ✘ | ✘ | ✔ |
| MultiAgentBench ( Zhu et al., 2025 ) | ✘ | ✘ | ✔ |
| SWE-Together ( Wu et al., 2026 ) | ✔ | ✘ | ✘ |
| ICAE-Bench ( Peng et al., 2026 ) | ✔ | ✘ | ✘ |
| SpecBench ( Hamblin et al., 2026 ) | ✔ | ✔ | ✘ |
| Domain | Tasks | Reqs. | L1 | L2 | L3 | Reqs./task |
|---|---|---|---|---|---|---|
| Coding | 90 | 539 | 90 | 164 | 285 | 5.99 |
| Math | 90 | 530 | 180 | 213 | 137 | 5.89 |
| QA | 90 | 482 | 90 | 97 | 295 | 5.36 |
| Total | 270 | 1,551 | 360 | 474 | 717 | 5.74 |
| Method | HumanEval | HotpotQA | DROP | GSM8K | MBPP | ProgSpec |
|---|---|---|---|---|---|---|
| GPT-4o-mini | 87.0 | 68.1 | 68.3 | 92.7 | 71.8 | 38.3 |
| MetaGPT ( Hong et al., 2024 ) | 87.0 | 67.3 | 77.3 | 82.8 | 66.6 | 8.9 |
| AgentCoder ( Huang et al., 2023 ) | 82.4 | 54.9 | 82.0 | 88.5 | 59.2 | 12.8 |
| G-Designer ( Zhang et al., 2024 ) | 88.6 | 68.1 | 71.5 | 95.1 | 71.3 | 11.1 |
| AFlow ( Zhang et al., 2025c ) | 94.7 | 73.5 | 80.6 | 93.5 | 83.4 | 39.9 |
| Meta-Agent ( Xu and Tai, 2026 ) | 96.0 ∗ | 69.5 ∗ | 82.7 ∗ | 93.7 ∗ | 84.6 ∗ | 42.9 |
| Setting | ProgSpec | HumanEval | HotpotQA |
|---|---|---|---|
| RepoMAS | 44.7 | 99.2 | 75.5 |
| w/o Structured Issue | 30.8 | 84.7 | 68.5 |
| w/o Patch Validation | 33.4 | 84.0 | 59.4 |
| w/o Structural Revision | 30.3 | 87.8 | 56.4 |
| w/o Local Re-execution | 32.4 | 84.7 | 69.7 |
| Issue Representation | ProgSpec | HumanEval | HotpotQA | Avg. |
|---|---|---|---|---|
| Direct Revision | 30.2 | 90.8 | 68.5 | 63.2 |
| Free-form Issue | 30.8 | 84.7 | 68.5 | 61.4 |
| Structured Issue w/o Evidence | 27.5 | 90.8 | 69.0 | 62.5 |
| Full Structured Issue | 44.7 | 99.2 | 75.5 | 73.1 |
| Metric | Value |
|---|---|
| Validation Precision | 62.5% |
| Validation Recall | 86.2% |
| Accepted-patch Error Rate | 37.5% |
| Regression Rate | 22.5% |
| Good patches: initial | 13.8% |
| Good patches: revised | 69.0% |
| Variant | ProgSpec | HumanEval | HotpotQA |
|---|---|---|---|
| Full RepoMAS | 44.7 | 99.2 | 75.5 |
| w/o Task Graph Revision | 25.8 | 90.1 | 55.8 |
| w/o Agent Graph Revision | 25.8 | 88.5 | 56.9 |
| w/o Structural Revision | 30.3 | 87.8 | 56.4 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Level | Example and evidence |
|---|---|
| L1 (Explicit) | The user asks the function to return a JSON object. This requirement is stated directly in the request. |
| L2 (Implicit) | The task asks for the real-valued expression to be evaluated. The condition follows directly from the domain of the real logarithm, although it is not stated in the request. |
| L3 (Latent) | The request does not specify how empty input should be handled. Inspection of an available repository test shows that empty input must return an empty result rather than raise an exception. |
| Backbone | Method | HumanEval | HotpotQA | DROP | GSM8K | MBPP | ProgSpec |
|---|---|---|---|---|---|---|---|
| GPT-5.5 | GPT-5.5 | 99.24 | 59.38 | 39.12 | 93.46 | 83.28 | 92.08 |
| MetaGPT ( Hong et al., 2024 ) | 98.47 | 32.88 | 3.38 | 71.18 | 81.23 | 89.95 | |
| AgentCoder ( Huang et al., 2023 ) | 99.24 | 50.62 | 33.75 | 92.61 | 83.58 | 93.02 | |
| G-Designer ( Zhang et al., 2024 ) | 100.00 | 57.00 | 36.50 | 93.36 | 84.46 | 94.20 | |
| AFlow ( Zhang et al., 2025c ) | 99.24 | 57.75 | 39.12 | 94.22 | 83.58 | 93.34 | |
| Meta-Agent ( Xu and Tai, 2026 ) | 98.47 | 54.37 | 25.50 | 92.89 | 80.06 | 89.50 |
| Metric | Score | 95% CI | Evaluation Unit |
|---|---|---|---|
| Annotation–review agreement | 91.4% ( ) | [88.0, 93.9] | 329 / 360 judgments |
| Requirement validity | 95.4% | [93.2, 96.9] | 477 / 500 requirements |
| Requirement independence | 98.2% | [96.6, 99.1] | 491 / 500 requirements |
| Evidence sufficiency | 97.4% | [95.6, 98.5] | 487 / 500 requirements |
| Ambiguity rate | 0.2% | [0.0, 1.1] | 1 / 500 requirements |
| Benchmark | Temperature | Max Tokens | Max Iterations |
|---|---|---|---|
| HumanEval | 0.2 | 4096 | 3 |
| HotpotQA / DROP / GSM8K / MBPP | 0.2 | 4096 | 3 |
| ProgSpec–Coding | 0.0 | 4096 | 3 |
| ProgSpec–Math / QA | 0.0 | 4096 | 3 |