ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
Organizations: Fudan University · University of Edinburgh · Chinese University of Hong Kong · Vivo AI · Huazhong University of Science and Technology · Shanghai Innovation Institute
Abstract
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
Figures & tables
| ToolBench | API-Bank | ||||||||||
| I1 | I2 | I3 | Avg | ||||||||
| Action | Backbone | Suc. | Path | Suc. | Path | Suc. | Path | Suc. | Path | Suc. | Acc. |
| Exploration-then-Execution | |||||||||||
| ReAct | GPT-4.1 | 43.2 | 28.9 0.4 | 29.6 | 17.3 1.2 | 29.5 | 10.2 2.4 | 34.1 | 18.8 0.3 | 20.0 | 44.3 1.1 |
| [2pt/2pt] DFS | GPT-4.1 | 45.1 | 29.2 0.3 | 28.7 | 17.4 1.3 | 27.9 | 9.6 2.1 | 33.9 | 18.7 0.2 | 20.0 | 45.4 0.5 |
| T-LLaMA 7B | 15.6 | 32.7 0.4 | 6.5 | 21.0 0.1 | 11.5 | 12.5 0.6 | 11.2 | 22.1 0.3 | 2.0 | 21.4 0.0 | |
| ToolBench | API-Bank | |||
|---|---|---|---|---|
| Method | Suc. | Path | Suc. | Acc. |
| Backbone | 22.13 | 23.49 | 16.00 | 35.88 |
| + SFT | 31.90 | 37.54 | 20.00 | 43.51 |
| [2pt/2pt] + SFT & RL(GRPO) | 47.02 | 46.51 | 14.00 | 52.67 |
| + SFT & RL (Ours) | 49.66 | 42.54 | 24.00 | 58.02 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Name | API lib. | Raw pairs | Cleaned pairs | # API | Category | Source |
| API calling | ||||||
| ToolBench ( Qin et al., 2024 ) | 37,366 | 149,693 | 105,877 | 9,483 | 50 | RapidAPI |
| APIGen ( Liu et al., 2024 ) | 25 | 20,728 | 9,962 | 25 | 10 | RapidAPI, OpenAPI Hub |
| ToolACE ( Liu et al., 2025 ) | 2,562 | 1,353 | 445 | 1,260 | 49 | Public API hubs |
| Simia ( Li et al., 2025d ) | 26 | 310,870 | 10,000 | 26 | 2 | Public API hubs |
| MCP server | ||||||
| Hyperparameter | Value |
| Max rounds per task ( ) | 12 |
| Retrieval top- per capability slot | 3 |
| Simulator error-injection rate | 5% |
| LLM decoding temperature (all agents) | 0.7 |
| Max tokens per agent call | 2,048 |
| Symbol | Scope | Value | Role |
| Eq. 7 | trajectory-advantage weight | ||
| Eq. 7 | phase-advantage weight | ||
| Eq. 7 | step-advantage weight | ||
| L1 retrieval grounding | |||
| L2 dependency declaration | |||
| L3 layered plan |
| State | Condition | Credited behavior | |
|---|---|---|---|
| Need | no controller, or DAG exhausted | emit a valid plan | , else |
| May | new observation or recoverable failure | no plan, or a valid new plan | , else |
| No | no new information | no plan | ; repeated plan: |
| Answer | final answer step | no <plan> |
| Q/F/U pattern | , , | , | , | |
|---|---|---|---|---|
| Meaning | insufficient evidence; not both complete and faithful | sufficient evidence, but unfaithful claims | faithful, but incomplete or under-evidenced | complete, faithful, and well-evidenced |
| Judge model | Quality | Faithfulness | Utility |
|---|---|---|---|
| Qwen3.6-27B | 89.0% | 76.0% | 78.0% |
| Qwen3-235B-A22B-Instruct-2507 (deployed) | 89.0% | 83.0% | 81.0% |