cs.LGOct 7, 2026

CERO: Where and When to Allocate Rollouts for RL Post-Training

Authors: Yiming Zong, Yige Wang, Xing Hu, Jiashuo Jiang, Zuo-Jun Max Shen

Organizations: Department of Industrial Engineering & Decision Analytics, Hong Kong University of Science and Technology · Innovation and Information Management, The University of Hong Kong · Faculty of Engineering & Faculty of Business and Economics, The University of Hong Kong

Abstract

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

    Jun 4, 2026Yiming Zong, Yige Wang, Jiashuo Jiang

  2. Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

    May 26, 2026Woojeong Kim, Ziyi Yang, Jing Nathan Yan +1RolloutToken Budget Allocation

  3. Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

    Jul 28, 2026Pixel Nomand, Elena Voss, Marcus Hale +1Reinforcement Learning With Verifiable RewardToken Budget Allocation