AgentFly: Scaling Agentic Reinforcement Learning with Unified Resource System
Organizations: Mohamed bin Zayed University of Artificial Intelligence
Abstract
Methods to build LLM agents have evolved from prompt engineering and supervised finetuning to agentic reinforcement learning (agentic RL). However, agentic RL remains bottlenecked by its surrounding systems: agents must interact with heterogeneous environments, such as sandboxes, model services, and external APIs. Their allocation, reuse, and lifecycle dominate rollout cost and cap the scale at which training becomes practical. In this work, we present AgentFly, an agentic RL framework built with a unified resource layer that treats each of these environments as a distinct, typed resource scheduled through one engine, with per-tool acquisition for multi-turn reuse, asynchronous backpressure, and rollout versus global-scoped lifecycles. AgentFly adopts a four-layer design: (I) agent layer that abstracts the agent, tool, and reward concepts, decomposing agentic RL into defining agents, tools, and reward functions; (II) rollout layer that composes these into agent loops and computes rewards; (III) context layer that organizes rollouts, injects contextual information, and arranges resources; and (IV) a low-level resource layer that performs resource management. We provide a suite of prebuilt tools and environments, demonstrate successful agent training across multiple tasks and models, and report the first controlled cross-framework throughput comparison against agentic RL frameworks.
Figures & tables
| Framework | Agent Abs. a | Async | Resource Scaling b | Multi-Modal c | Backends |
| AgentFly | ● | ● | ● | T+I | verl |
| RAGEN | ○ | ○ | ◐ | ○ | verl |
| verl-agent | ○ | ○ | ◐ | T+I | verl |
| RL-Factory | ◐ | ● | ○ | T+I | verl |
| agent-lightning | ◐ | ● | ○ | N/A | multiple d |
| verl-tool | ◐ | ● | ○ | T+I+V | verl |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Algorithm | Initial (%) | Final (%) | Best (%) | Steps to 80% | Time (h) |
|---|---|---|---|---|---|
| PPO | 31.4 | 95.2 | 97.1 | 50 | 12.5 |
| REINFORCE++ | 30.7 | 92.6 | 94.3 | 70 | 12.9 |
| RLOO | 30.0 | 91.9 | 94.3 | 80 | 14.2 |
| GRPO | 26.4 | 88.1 | 90.0 | 100 | 16.2 |
| WebShop (in-distribution) | WebShop (held-out) | ALFWorld | ||||||
|---|---|---|---|---|---|---|---|---|
| Rollout | Algorithm | 1.5B | 3B | 7B | 1.5B | 3B | 7B | 1.5B |
| Step | GRPO | 58.9 | 50.1 | 65.8 | 53.1 | 39.8 | 48.0 | 75.5 |
| Step | GiGPO | 65.3 | 66.0 | 73.4 | 52.1 | 46.7 | 59.0 | 89.3 |
| Chain | GRPO | 64.7 | 69.2 | 68.0 | 52.3 | 52.5 | 58.0 | 91.0 |
| Chain | GiGPO | 66.3 | 70.1 | 73.7 | 52.5 | 55.3 | 61.1 | 98.1 |
| SearchQA | WebShop | SimuScene | R2E-Gym | |
| Model | Qwen3.5-{4B, 9B}-Base | Qwen3.5-{4B, 9B} | DeepSeek-R1-Distill-Qwen-7B (SFT) | Qwen3-32B |
| Resource | retrieval API | shared container | VLM judge API | container per task |
| Training data | NQ + HotpotQA | WebShop (small) | SimuScene | R2E-Gym-Lite |
| Reward | exact match | WebShop score | VLM-judged pass | unit tests |
| Tasks group size | ||||
| Max turns | 3 | 15 | 1 | 30 |
| AgentFly | verl-agent | RAGEN | |
| Model | Qwen2.5-3B-Instruct | ||
| Algorithm | GRPO | ||
| Tasks group size | |||
| Max turns | 15 | 15 | 15 (9 actions) |
| Max tokens per turn | 384 | 384 | 400 |
| History in context | full | last 2 turns | full |
| AgentFly | SkyRL | |
| Model | Qwen3-4B | |
| Algorithm | GRPO | |
| Tasks group size | ||
| Trajectories per step | 64 | |
| Max turns | 20 | |
| Max tokens per turn | 4,096 | |
| Token drift | WebShop | ALFWorld | |||
| WebShop | Step | Chain | Step | Chain | |
| Model | Qwen3.5-4B | Qwen2.5-{1.5B, 3B, 7B}-Instruct | Qwen2.5-1.5B-Instruct | ||
| Algorithm | GRPO | GRPO, GiGPO | GRPO, GiGPO | ||
| GiGPO normalization | – | mean | mean and std | ||
| Tasks group size | |||||
| Max turns | 15 | 15 | 15 | 50 | 50 |
| Model | Dev | Creative | CAD | Sci | Office | OS | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | Text | Icon | |
| QwenVL2.5-3B | 16.2 | 1.4 | 23.3 | 1.4 | 10.2 | 4.7 | 38.2 | 6.4 | 24.3 | 3.8 | 15.0 | 1.1 |
| QwenVL2.5-3B w/ GRPO | 42.9 | 2.1 | 37.4 | 7.7 | 17.8 | 1.6 | 49.3 | 9.1 | 55.4 | 20.8 | 39.3 | 7.9 |
| QwenVL2.5-7B | 33.1 | 2.1 | 23.7 | 3.5 | 12.2 | 6.3 | 36.8 | 7.3 | 37.8 | 7.5 | 30.8 | 6.9 |
| QwenVL2.5-7B w/ GRPO | 46.8 | 3.4 | 39.7 | 8.3 | 19.6 | 5.1 | 52.8 | 10.7 | 58.6 | 22.7 | 42.8 | 9.3 |
| Tool | Arguments | Behavior |
|---|---|---|
| read_file | path , start_line , end_line | Returns the file with line numbers, optionally restricted to a line range. |
| create_file | path , content | Creates a new file; fails if the path already exists. |
| edit_file | path , search_block , replace_block | Replaces an exact block of text; fails if the block is not found or matches more than one location. The previous content is saved for undo. |
| undo_edit | path | Restores the file to its content before the last edit, using a per-file snapshot stack stored in a hidden directory of the workspace. |
| list_files | path , max_depth | Lists files recursively (default depth 3), skipping hidden, cache, and build directories and binary files. |
| grep_search | pattern , path , include | Searches for a regular expression with grep -rn (with a Python fallback), optionally filtered by a file glob; returns at most 1,000 matches. |