Methods to build LLM agents have evolved from prompt engineering and supervised finetuning to agentic reinforcement learning (agentic RL). However, agentic RL remains bottlenecked by its surrounding systems: agents must interact with heterogeneous environments, such as sandboxes, model services, and external APIs. Their allocation, reuse, and lifecycle dominate rollout cost and cap the scale at which training becomes practical. In this work, we present AgentFly, an agentic RL framework built with a unified resource layer that treats each of these environments as a distinct, typed resource scheduled through one engine, with per-tool acquisition for multi-turn reuse, asynchronous backpressure, and rollout versus global-scoped lifecycles. AgentFly adopts a four-layer design: (I) agent layer that abstracts the agent, tool, and reward concepts, decomposing agentic RL into defining agents, tools, and reward functions; (II) rollout layer that composes these into agent loops and computes rewards; (III) context layer that organizes rollouts, injects contextual information, and arranges resources; and (IV) a low-level resource layer that performs resource management. We provide a suite of prebuilt tools and environments, demonstrate successful agent training across multiple tasks and models, and report the first controlled cross-framework throughput comparison against agentic RL frameworks.
Figures & tables
Figure 1 : Architecture of AgentFly. The agent layer holds the user-defined agent, tools, and reward; the rollout layer runs the multi-turn agent loop for many chains in parallel; the context layer injects rollout information and brokers resources; and the resource layer manages one pool per typed resource specification, launched by pluggable runners. Together, the context and resource layers form the unified resource system. Completed rollouts are sent to the RL trainer ( verl ) as trajectories, and updated weights are returned to the agent.
Figure 2 : Fully asynchronous rollout in AgentFly. Left: within a rollout, generation, tool calls, reward computation, and resource operations (peach) are each awaited asynchronously, so a slow step in one rollout does not block the others. Right: schematic schedule of three rollouts with two turns each: in a batch-synchronous design every phase waits for the slowest rollout (dotted), whereas in AgentFly each rollout proceeds independently and the batch finishes earlier. Durations are illustrative.
Figure 3 : One acquisition interface for two resource kinds. Ag( Orange lines declare the resource a function needs; Ag( green lines acquire it through the context layer. Left: a tool that runs shell commands in an isolated container whose pool holds at most 24 instances; acquiring under the rollout id makes all turns of a rollout reuse one container, released when the rollout ends. Right: a reward that judges the final answer with an external model API; the same call acquires a model resource with global scope, shared across rollouts ( judge_prompt builds the chat messages). Neither function contains allocation or cleanup logic.
Figure 4
Figure 6 : Training reward on four tasks that use different kinds of resources: a retrieval API (SearchQA), a container shared across rollouts (WebShop), a VLM judge API (SimuScene), and one container per task (R2E-Gym). Bold lines are exponential moving averages; faint lines are per-step values.
Figure 7 : Training throughput (trajectories/s) on WebShop and R2E-Gym, with all frameworks on 8 GPUs. End-to-end includes the policy update; rollout-only covers generation and environment interaction. Error bars show one standard deviation across steps.
Figure 8 : Training on sampled token IDs versus re-tokenized text on WebShop with Qwen3.5-4B (chain rollout, GRPO). The two runs differ only in whether the sampled token IDs are kept; the re-tokenized setting is shown for two runs. KL loss and gradient norm are exponential moving averages over per-step values (faint), on a log scale.
Framework
Agent Abs. a
Async
Resource Scaling b
Multi-Modal c
Backends
AgentFly
●
●
●
T+I
verl
RAGEN
○
○
◐
○
verl
verl-agent
○
○
◐
T+I
verl
RL-Factory
◐
●
○
T+I
verl
agent-lightning
◐
●
○
N/A
multiple d
verl-tool
◐
●
○
T+I+V
verl
Table 1 : Comparison of agentic RL frameworks. ● full / native, ◐ partial or depends on configuration, ○ none. a Agent abstraction including agent, tool, and reward abstractions. b ◐ indicates inherited from underlying infrastructure (e.g., Ray, SkyRL-Agent on K8s); ● indicates dedicated resource system (e.g., AgentFly’s resource engine). c T = text, I = image, V = video. d LangChain, AutoGen, CrewAI, OpenAI SDK for agent rollout. e SkyRL-train, verl, Tinker. f slime (SGLang for rollout, Megatron-LM for training).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Algorithm
Initial (%)
Final (%)
Best (%)
Steps to 80%
Time (h)
PPO
31.4
95.2
97.1
50
12.5
REINFORCE++
30.7
92.6
94.3
70
12.9
RLOO
30.0
91.9
94.3
80
14.2
GRPO
26.4
88.1
90.0
100
16.2
Appendix
Table 2 : ALFWorld results of Qwen2.5-3B-Instruct with different RL algorithms. Success is measured on valid_seen ; Final is the mean of the last three evaluations (steps 180–200), and Steps to 80% is the first evaluation at or above 80% success. Time is the training time of the 200-step run on 8 H200 GPUs.
WebShop (in-distribution)
WebShop (held-out)
ALFWorld
Rollout
Algorithm
1.5B
3B
7B
1.5B
3B
7B
1.5B
Step
GRPO
58.9
50.1
65.8
53.1
39.8
48.0
75.5
Step
GiGPO
65.3
66.0
73.4
52.1
46.7
59.0
89.3
Chain
GRPO
64.7
69.2
68.0
52.3
52.5
58.0
91.0
Chain
GiGPO
66.3
70.1
73.7
52.5
55.3
61.1
98.1
Appendix
Table 3 : Success rate (%) of chain and step rollout with GRPO and GiGPO. WebShop: in-distribution success on the 512 validation tasks of Feng et al. [4] (mean of the last three evaluations) and success on 512 held-out tasks whose products never appear in training (final checkpoint). ALFWorld: success on valid_seen (mean of the last three evaluations). All models are Qwen2.5-Instruct.
SearchQA
WebShop
SimuScene
R2E-Gym
Model
Qwen3.5-{4B, 9B}-Base
Qwen3.5-{4B, 9B}
DeepSeek-R1-Distill-Qwen-7B (SFT)
Qwen3-32B
Resource
retrieval API
shared container
VLM judge API
container per task
Training data
NQ + HotpotQA
WebShop (small)
SimuScene
R2E-Gym-Lite
Reward
exact match
WebShop score
VLM-judged pass
unit tests
Tasks × group size
256×5
32×8
32×8
16×8
Max turns
3
15
1
30
Appendix
Table 4: Settings of the tasks with different resources ( Figure 6 ). Tasks × group size is the number of trajectories per step.
AgentFly
verl-agent
RAGEN
Model
Qwen2.5-3B-Instruct
Algorithm
GRPO
Tasks × group size
32×8
Max turns
15
15
15 (9 actions)
Max tokens per turn
384
384
400
History in context
full
last 2 turns
full
Appendix
Table 5: Settings of the WebShop throughput comparison.
AgentFly
SkyRL
Model
Qwen3-4B
Algorithm
GRPO
Tasks × group size
8×8
16×4
Trajectories per step
64
Max turns
20
Max tokens per turn
4,096
Appendix
Table 6: Settings of the R2E-Gym throughput comparison.
Token drift
WebShop
ALFWorld
WebShop
Step
Chain
Step
Chain
Model
Qwen3.5-4B
Qwen2.5-{1.5B, 3B, 7B}-Instruct
Qwen2.5-1.5B-Instruct
Algorithm
GRPO
GRPO, GiGPO
GRPO, GiGPO
GiGPO normalization
–
mean
mean and std
Tasks × group size
16×8
Max turns
15
15
15
50
50
Appendix
Table 7: Settings of the token-drift runs ( Section 4.3 ) and the chain and step rollout runs ( Section 4.4 ). Mini-batches count whole trajectories for chain rollout and single-step samples for step rollout.
Figure 9 : Training reward of Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on four tasks. Bold lines are exponential moving averages; faint lines are per-step values.
Model
Dev
Creative
CAD
Sci
Office
OS
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
Text
Icon
QwenVL2.5-3B
16.2
1.4
23.3
1.4
10.2
4.7
38.2
6.4
24.3
3.8
15.0
1.1
QwenVL2.5-3B w/ GRPO
42.9
2.1
37.4
7.7
17.8
1.6
49.3
9.1
55.4
20.8
39.3
7.9
QwenVL2.5-7B
33.1
2.1
23.7
3.5
12.2
6.3
36.8
7.3
37.8
7.5
30.8
6.9
QwenVL2.5-7B w/ GRPO
46.8
3.4
39.7
8.3
19.6
5.1
52.8
10.7
58.6
22.7
42.8
9.3
Appendix
Table 8 : Performance of ScreenSpot-Pro GUI grounding.
Tool
Arguments
Behavior
read_file
path , start_line , end_line
Returns the file with line numbers, optionally restricted to a line range.
create_file
path , content
Creates a new file; fails if the path already exists.
edit_file
path , search_block , replace_block
Replaces an exact block of text; fails if the block is not found or matches more than one location. The previous content is saved for undo.
undo_edit
path
Restores the file to its content before the last edit, using a per-file snapshot stack stored in a hidden directory of the workspace.
list_files
path , max_depth
Lists files recursively (default depth 3), skipping hidden, cache, and build directories and binary files.
grep_search
pattern , path , include
Searches for a regular expression with grep -rn (with a Python fallback), optionally filtered by a file glob; returns at most 1,000 matches.
Appendix
Table 9: File tools available to SWE agents. All tools run inside the rollout’s container.