Organizations: Department of Industrial Engineering & Decision Analytics, Hong Kong University of Science and Technology · Innovation and Information Management, The University of Hong Kong · Faculty of Engineering & Faculty of Business and Economics, The University of Hong Kong
Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.
Figures & tables
Figure 1 : Schematic comparison of rollout optimization. (a) Typical adaptive methods redistribute a fixed per-round budget across prompts. (b) CERO coordinates prompt admission and per-round budget pacing under a shared finite training budget, with admitted prompts receiving fixed-size response groups.
Base Model
Method
AIME
AMC
MATH
MINERVA
Olympiad
Avg.
DSR1-Q1.5B
GRPO
5.62
32.23
66.12
21.90
25.07
30.19
KnapsackRL
4.58
32.38
65.14
20.93
24.59
29.53
CERO (ours)
7.92
38.03
68.05
22.66
27.41
32.81
Random-matched CERO
4.17
31.02
64.54
21.48
24.69
29.18
Qwen3-4B
GRPO
13.54
50.98
80.10
30.72
40.94
43.26
KnapsackRL
17.50
53.24
80.08
31.36
41.26
44.69
Table 1 : Evaluation performance (avg@16) comparison across different models and benchmarks.
Figure 3Figure 4
Method
Avg.
CERO
32.26 ± 0.94
CERO-fixed
31.13 ± 0.72
CERO-greedy
30.97 ± 0.38
CERO-preset
30.71 ± 0.54
Table 2 : Prompt-selection and pacing ablations on DSR1-Q1.5B. Avg. denotes the five-benchmark avg@16 macro-average (mean ± std over 3 seeds).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Base Model
Method
AIME
AMC
MATH
MINERVA
Olympiad
Avg.
DSR1-Q1.5B
GRPO
20.00
55.42
83.40
46.69
40.89
49.28
KnapsackRL
20.00
56.63
82.80
42.65
40.59
48.53
CERO (ours)
20.00
62.65
84.20
44.85
43.26
50.99
Random-matched CERO
13.33
51.81
81.40
45.96
39.70
46.44
Qwen3-4B
GRPO
40.00
79.52
92.00
51.10
62.67
65.06
KnapsackRL
46.67
80.72
92.20
51.84
64.00
67.09
Appendix
Table 3 : Evaluation performance (pass@16) comparison across different models and benchmarks.
Figure 6 : Calibration of CERO’s reward-variation proxy on DSR1-Q1.5B. Points show predicted and observed bin means across seven prediction bins; vertical bars indicate nominal 95% within-run intervals for the observed means. The dashed diagonal marks perfect calibration.
Figure 7 : Empirical CDF of normalized terminal prompt exposure zi,K/R for CERO and GRPO on DSR1-Q1.5B. The dashed line marks the equal-share exposure R=B/M .
Figure 8 : Training-compute comparison on DeepSeek-R1-Distill-Qwen-1.5B under a matched response budget. The left panel reports end-to-end GPU-hours, while the right panel reports core training time normalized by generated response tokens.
LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt, despite large differences in the training signal different prompts provide. In this paper, we study adaptive rollout allocation under a fixed global budget and formulate the problem as online resource allocation with prompt-level diminishing returns. Our method, CERO, maintains a Beta posterior over each prompt's success probability and uses the posterior expected Bernoulli variance as a Bayesian estimate of the value of additional rollouts. We use this estimate to construct a concave, saturating utility over cumulative allocations, yielding an objective in which decisions across prompts and epochs are coupled by the global budget. Since the resulting objective is temporally nonseparable, we derive a Fenchel-dual reformulation and update both prompt-level and budget-level dual variables via projected online gradient descent. Under fixed prompt utilities, we prove an O(K) regret bound against the offline allocation benchmark. Experiments on mathematical-reasoning problems show that CERO consistently outperforms GRPO across multiple open-weight LLMs and benchmarks, demonstrating that adaptive rollout budgeting can improve sample efficiency.
Yiming Zong, Yige Wang, Jiashuo Jiang
Department of Industrial Engineering & Decision Analytics, Hong Kong University of Science and Technology
Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout generation dominates the computational cost of training. Group-based policy optimization methods compute advantages from multiple rollouts per prompt, yet they indiscriminately allocate budget to prompts with collapsed reward distributions, wasting expensive rollouts on negligible learning signals. We demonstrate that group-based updates are most effective in regimes of high reward variance. Since the policy evolves throughout training, prompt informativeness must be estimated online rather than precomputed, but exhaustively evaluating every prompt is computationally prohibitive. We introduce Pilot-Commit, a budget-aware rollout allocation framework for group-based RL post-training. Pilot-Commit decouples prompt evaluation from exploitation: a pilot stage estimates per-prompt informativeness using a fraction of the budget, and the remaining rollouts are allocated to high-leverage prompts while low-signal prompts are skipped. Across multiple math reasoning benchmarks and model scales from 1.5B to 14B parameters, Pilot-Commit matches baseline accuracy with significantly lower sampling costs, reaching target accuracy up to 1.9× faster than GRPO and 4.0× faster than DAPO in cumulative rollouts.
Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal. Existing remedies either oversample a larger candidate pool and discard saturated prompts (dynamic sampling), paying heavy extra rollouts, or predict prompt difficulty before sampling, which is fragile under a shifting policy. We observe that a group's effectiveness is usually decided early, within the first few of its rollouts, so spending a full group on an already-decided prompt is wasteful. We cast per-step rollout collection as a budget-constrained sequential allocation (optimal stopping) problem and introduce SARA (Sequential Adaptive Rollout Allocation). SARA maintains a Beta posterior over each prompt's success rate, evaluates a closed-form predictor of group effectiveness, and applies a two-threshold, SPRT-style rule that commits effective groups, abandons saturated ones after a short probe, and reallocates the freed budget to fresh prompts, without any extra prediction rollouts. We prove abandonment reliability, expected rollout savings, fixed-budget yield dominance, and a link between effective-group yield and the GRPO gradient norm. On mathematical reasoning and planning with 1.5B/3B models on a single GPU, SARA matches DPS (both below the DS oracle) while using 22% fewer rollouts than DS; composing SARA with DPS yields the best accuracy, slightly above DS, at 67% fewer rollouts (near-uniform cost).
Pixel Nomand, Elena Voss, Marcus Hale +1
University of Wisconsin–Madison · University of Washington