SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
Organizations: Harbin Institute of Technology, Shenzhen · The Chinese University of Hong Kong, Shenzhen · Nanyang Technological University
Abstract
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
Figures & tables
| Method | Pass@1 | Pass@32 | Pass@128 | Pass@256 |
| Base model | 6.80 | 53.85 | 76.32 | 84.37 |
| GRPO ( Shao et al., 2024 ) | 30.46 | 45.80 | 49.28 | 50.62 |
| PKPO (T=16) ( Walder and Karkhanis, 2025 ) | 19.93 | 70.32 | 84.92 | 90.16 |
| MaxRL (token-mean) ( Tajwar et al., 2026 ) | 32.60 | 59.41 | 68.44 | 72.21 |
| MaxRL (Seqnorm) ( Tajwar et al., 2026 ) | 32.08 | 56.66 | 64.58 | 68.24 |
| DARS-HW † ( Yang et al., 2025b ) | 34.25 | 52.38 | 57.31 | 59.06 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | ||||||||||||
| – | Direct CE | 60.81 | 70.26 | 77.38 | 82.78 | 86.95 | 90.18 | 92.68 | 94.59 | 96.05 | 97.16 | 98.00 |
| MaxRL | 4.43 | 4.97 | 5.37 | 5.71 | 6.05 | 6.47 | 7.13 | 8.24 | 10.21 | 13.67 | 19.35 | |
| SERA | 12.37 | 13.61 | 14.54 | 15.26 | 15.82 | 16.25 | 16.59 | 16.86 | 17.09 | 17.29 | 17.49 | |
| MaxRL | 36.59 | 41.30 | 44.79 | 47.43 | 49.45 | 51.01 | 52.24 | 53.22 | 54.00 | 54.62 | 55.14 | |
| SERA | 54.80 | 63.53 | 70.26 | 75.46 | 79.52 | 82.75 | 85.37 | 87.51 | 89.25 | 90.65 | 91.79 |
| Estimator | MAE | Pearson | |
| Pilot estimate | 0.0634 | 0.9879 | |
| Historical estimate | 0.0386 | 0.8134 |
| Normalization | Method | ||||||
| – | Base model | 1.27 | 2.49 | 4.75 | 8.83 | 15.50 | 25.17 |
| Token-mean | GRPO | 34.00 | 34.80 | 35.43 | 36.01 | 36.61 | 37.28 |
| Token-mean | MaxRL | 52.95 | 59.16 | 63.16 | 66.16 | 68.62 | 70.75 |
| Token-mean | SERA | 49.98 | 58.19 | 63.85 | 68.06 | 71.49 | 74.45 |
| Seqnorm | MaxRL | 52.80 | 59.42 | 63.74 | 66.96 | 69.63 | 71.94 |
| Seqnorm | SERA | 55.61 | 61.66 | 65.93 | 69.45 | 72.43 | 74.97 |
| Pass@32 | Pass@128 | Pass@256 | |
| CI |
| Dataset | Method | ||||||||||
| BeyondAIME | Base | 0.73 | 1.41 | 2.61 | 4.57 | 7.36 | 10.96 | 15.49 | 21.20 | 28.13 | 35.00 |
| MaxRL | 3.29 | 5.25 | 7.71 | 10.82 | 14.82 | 19.53 | 24.78 | 30.53 | 36.53 | 42.00 | |
| SERA | 3.07 | 5.06 | 7.64 | 10.82 | 14.76 | 19.55 | 25.36 | 32.26 | 40.14 | 49.00 | |
| AIME 2025 | Base | 3.26 | 6.01 | 10.36 | 16.17 | 22.59 | 29.21 | 36.42 | 44.21 | 53.12 | 63.33 |
| MaxRL | 10.05 | 15.64 | 21.65 | 27.37 | 33.02 | 38.86 | 44.49 | 49.91 | 56.26 | 63.33 | |
| SERA | 10.18 | 16.28 | 23.04 | 29.19 | 34.77 | 40.46 | 46.30 | 52.85 | 61.26 | 70.00 |
| Dataset | Pass@32 | Pass@128 | Pass@256 | Pass@512 |
| BeyondAIME | ||||
| AIME 2025 | ||||
| MATH-500 | ||||
| OlympiadBench |
| Dataset | Method | ||||||||||
| BeyondAIME | Base | 4.02 | 6.16 | 8.51 | 11.08 | 14.33 | 18.61 | 23.90 | 30.35 | 37.75 | 45.00 |
| MaxRL | 6.97 | 10.53 | 14.46 | 18.62 | 23.28 | 28.50 | 34.08 | 40.27 | 47.50 | 55.00 | |
| SERA | 7.78 | 11.34 | 15.78 | 21.08 | 27.17 | 34.01 | 41.39 | 48.81 | 55.91 | 63.00 | |
| AIME 2025 | Base | 6.82 | 11.78 | 18.33 | 25.23 | 31.48 | 37.23 | 43.08 | 49.90 | 57.41 | 63.33 |
| MaxRL | 16.37 | 22.75 | 28.50 | 33.84 | 39.43 | 45.34 | 51.72 | 59.17 | 66.77 | 73.33 | |
| SERA | 18.37 | 23.85 | 29.30 | 35.38 | 41.88 | 48.40 | 54.97 | 61.20 | 67.61 | 73.33 |
| Dataset | Pass@32 | Pass@128 | Pass@256 | Pass@512 |
| BeyondAIME | ||||
| AIME 2025 | ||||
| MATH-500 | ||||
| Minerva Math |