CERO: Where and When to Allocate Rollouts for RL Post-Training
Organizations: Department of Industrial Engineering & Decision Analytics, Hong Kong University of Science and Technology · Innovation and Information Management, The University of Hong Kong · Faculty of Engineering & Faculty of Business and Economics, The University of Hong Kong
Abstract
Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.
Figures & tables
| Base Model | Method | AIME | AMC | MATH | MINERVA | Olympiad | Avg. |
| DSR1-Q1.5B | GRPO | 5.62 | 32.23 | 66.12 | 21.90 | 25.07 | 30.19 |
| KnapsackRL | 4.58 | 32.38 | 65.14 | 20.93 | 24.59 | 29.53 | |
| CERO (ours) | 7.92 | 38.03 | 68.05 | 22.66 | 27.41 | 32.81 | |
| Random-matched CERO | 4.17 | 31.02 | 64.54 | 21.48 | 24.69 | 29.18 | |
| Qwen3-4B | GRPO | 13.54 | 50.98 | 80.10 | 30.72 | 40.94 | 43.26 |
| KnapsackRL | 17.50 | 53.24 | 80.08 | 31.36 | 41.26 | 44.69 |
| Method | Avg. |
| CERO | 32.26 0.94 |
| CERO-fixed | 31.13 0.72 |
| CERO-greedy | 30.97 0.38 |
| CERO-preset | 30.71 0.54 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Base Model | Method | AIME | AMC | MATH | MINERVA | Olympiad | Avg. |
| DSR1-Q1.5B | GRPO | 20.00 | 55.42 | 83.40 | 46.69 | 40.89 | 49.28 |
| KnapsackRL | 20.00 | 56.63 | 82.80 | 42.65 | 40.59 | 48.53 | |
| CERO (ours) | 20.00 | 62.65 | 84.20 | 44.85 | 43.26 | 50.99 | |
| Random-matched CERO | 13.33 | 51.81 | 81.40 | 45.96 | 39.70 | 46.44 | |
| Qwen3-4B | GRPO | 40.00 | 79.52 | 92.00 | 51.10 | 62.67 | 65.06 |
| KnapsackRL | 46.67 | 80.72 | 92.20 | 51.84 | 64.00 | 67.09 |