Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.
Figures & tables
Figure 1: Overview of RPG in MLE , where our goal is to optimize for reward per unit of time in agentic RL contexts . Our agent under policy πθ is asked to self-improve upon its previous MLE scripts from buffer B , and each action a is a proposed script that executes in variable time δ and admits a performance score r such as AUC. ρ^ is estimated from past samples in buffer B and captures the cost per unit of time. Relative reward r~ is then used in PPO to perform RL training. Maximizing the relative reward optimizes for reward rate, giving higher reward per unit of time and more efficient RL agents.
Algorithm 1 RPG
Figure 2: Our proposed approach NPG - NIW obtains similar or lower regret than C-UCB while avoiding exhaustive search over policies . Standard policy gradient (SPG- NIW ) also performs worse than NPG , verifying our theory that NPG ’s decoupling helps reward-rate RL . y-axis is regret at the final step over 100 trials, x-axis is the number of arms K , and 4 environments are considered: (E1) independent reward and time, (E2) time and reward have 0.8 correlation, (E3) time and reward have −0.8 correlation, and (E4) reward and time are lognormal with +0.5 correlation.
Figure 4
Figure 4: Budget curve and an example self-improvement step on NanoGPT . (Left) At the final budget, RPG improves over vanilla by 85.7% . (Right) Example of a self-improvement edit by RPG , which rewrites Newton–Schulz in the Muon optimizer as a triton kernel with a compiled iteration and adds a learning-rate warmup. The new edit crosses the loss target in less time.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Example prompt on MLE-Bench .
Figure 6: Example prompt for self-improvement . previous_plan_code is filled with the executed code of the selected transition, which is the action a . previous_plan_error contains the grader’s own message.
Group
Setting
Value
Model
actor and reference
Qwen3.5-4B
critic
Qwen3.5-4B, token-classification head
rollout engine
vLLM, bfloat16
attention, checkpointing
sdpa , on
PPO
advantage estimator
GAE, γ=1.0 , λ=1.0
actor / critic learning rate
1×10−6 / 1×10−5
Appendix
Table 2: MLE-Bench configuration . The upper block is shared by RPG and vanilla. The lower block is where they differ.
Figure 9: Example action for NanoGPT
Group
Setting
Value
Model
actor and reference
Qwen3.5-4B
critic
Qwen3.5-4B, token-classification head
rollout engine
vLLM, bfloat16
PPO
advantage estimator
GAE, γ=1.0 , λ=1.0
actor / critic learning rate
1×10−6 / 1×10−5
train batch = mini-batch
64
Appendix
Table 3: NanoGPT configuration . The upper block is shared by RPG and vanilla; the lower block is the complete set of settings on which they differ.
Figure 10: Regret against step for ∣X∣=4 contexts .
Figure 11: Tuning grid for C-UCB . The horizontal axis is c2 and the vertical axis is c1 .
Figure 12: Tuning grid for the NIW - NPG , over the learning rate and accumulation window.
Figure 13: Tuning grid for the SPG-NIW , over the same learning rate and accumulation window grid as Figure 12 .
Figure 14: Regret at the horizon against the number of arms K , one panel per environment family, for the approximate C-UCB setting.
Figure 15: Regret against step for the large context setting , with environment families down the rows and arm counts across the columns.
Figure 16: Tuning grid for approximate C-UCB . The horizontal axis is c2 and vertical axis is c1 .
Figure 17: Tuning grid for the NIW - NPG , over the learning rate and accumulation window.
Figure 18: Tuning grid for the SPG-NIW , over the same learning rate and accumulation window grid as Figure 17 .
Figure 19: Per-task learning curves on MLE-Bench Lite , the evaluation method is same as Table 1 . Table 1 reports each task at the final time budget, the figure shows the entire learning process.
Figure 20: Mean cumulative time (x-axis) and mean cumulative reward (y-axis). RPG optimizes for long-term reward per unit of time, so its curve goes higher as mean cumulative time increases.
Figure 21: Budget curve and an example self-improvement step on NanoGPT, with root script included for reference . (Left) At the final budget, RPG improves over vanilla by 85.7% . Time 0 reflects the root parent script’s validation loss. We do not count the root script’s loss in the best crossings to reflect the dynamics of the policies. (Right) Example of a self-improvement edit by RPG , which rewrites Newton–Schulz in the Muon optimizer as a triton kernel with a compiled iteration and adds a learning-rate warmup. The new edit crosses the loss target using shorter time.
Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the rollout stage generates trajectories while invoking tools, producing long-tailed and non-stationary workloads that expose two fundamental challenges. First, due to the long-tailed response distribution, a small fraction of trajectories dominates rollout makespan.Second, rollout and training differ in their compute patterns, memory demands, and sensitivity to sequence length. As the policy evolves, shifts in the workload distribution further change their relative resource demands, making it difficult to maintain balanced execution across the two stages. We present Libra, an adaptive runtime for agentic RL post-training with two complementary components: (1) intra-stage scheduling via a Causality-Guided Bucket Scheduler that routes requests across execution buckets with different parallelism configurations, reducing delays from rollout stragglers; and (2) cross-stage coordination that dynamically reallocates workers between rollout and training as the workload changes. It moves workers between the two stages through a non-blocking protocol without interrupting ongoing training. Evaluated on a 48x NVIDIA A800 GPU cluster and a 160x Ascend 910B3 NPU cluster across three agentic benchmarks, Libra achieves up to 4.2x higher throughput and up to 2.7x faster reward convergence
Kaiwen Chen, Xin Tan, Jingzong Li +6
The Chinese University of Hong Kong · The Hang Seng University of Hong Kong · Independent Researcher
Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage energy under Group Relative Policy Optimization (GRPO), and success-failure pair count. We propose Prefix Sampling (PS), which replays self-generated trajectory prefixes to steer skewed groups toward this regime: successful prefixes give mostly failing groups a head start, while failing prefixes handicap mostly passing groups. Replayed states are reconstructed through the existing rollout path, and replayed tokens are masked from the loss so optimization applies only to current-policy continuations. On SWE-bench Verified, PS reaches the baseline high-score regime within evaluation variability while delivering 2.01x and 1.55x end-to-end wall-clock speedups on Qwen3-14B and Qwen3-32B; the 14B peak improves from 0.274 to 0.295. AIME 2025 experiments on 4B and 8B show the same pass-rate-control pattern, and 4B ablations attribute gains to replay, bidirectional coverage, and adaptive control.
Agentic reinforcement learning trains large language models using multi-turn trajectories that interleave long reasoning traces with short environment-facing actions. Common policy-gradient methods, such as PPO and GRPO, treat each token in a trajectory equally, leading to uniform credit assignment. In this paper, we critically demonstrate that such uniform credit assignment largely misallocates token-level training signals. From an energy-based modeling perspective, we show that token-level training signals, quantified by their correlations with reward variance of different rollouts sampled from a given prompt, concentrate sharply on action tokens rather than reasoning tokens, even though action tokens account for only a small fraction of the trajectory. We refer to this phenomenon as the Action Bottleneck. Motivated by this observation, we propose an embarrassingly simple token reweighting approach, ActFocus, that downweights gradients on reasoning tokens, along with an additional energy-based redistribution mechanism that further increases the weights on action tokens with higher uncertainty. Across four environments and different model sizes, ActFocus consistently outperforms PPO and GRPO, yielding final-step gains of up to 65.2 and 63.7 percentage points, respectively, without any additional runtime or memory cost.
Langzhou He, Junyou Zhu, Yue Zhou +7
University of Illinois Chicago · Potsdam Institute for Climate Impact Research · Technical University of Berlin +3