Does Scaling Reinforcement Learning Really Require More Training?
Organizations: University of Illinois at Urbana-Champaign · Tsinghua University
Abstract
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
Figures & tables
| Policy | Metric | AIME24 | AIME25 | HMMT25 | BRUMO25 | CMIMC25 | AMC23 | MATH | Minerva | Olympiad | Avg. |
| DeepSeek | |||||||||||
| Donor | Acc. | 47.50 | 35.83 | 19.17 | 46.67 | 21.88 | 88.13 | 90.75 | 50.64 | 64.80 | 51.71 |
| Tokens | 7,651 | 6,912 | 7,574 | 7,194 | 8,525 | 4,753 | 3,115 | 4,582 | 5,570 | 6,208 | |
| Anchor | Acc. | 50.83 | 36.67 | 20.00 | 49.17 | 22.50 | 90.00 | 91.60 | 52.02 | 66.58 | 53.26 |
| Tokens | 8,065 | 7,927 | 8,109 | 7,892 | 9,147 | 5,472 | 3,656 | 4,796 | 6,246 | 6,812 | |
| Surge | Acc. | 54.17 | 38.33 | 23.33 | 51.67 | 25.63 | 92.50 | 92.40 | 52.76 | 67.32 | 55.35 |
| Task. Implement same_chars(s0, s1) : compare the two strings’ character sets. Order and repetition do not matter; case matters. | ||
|---|---|---|
| Donor | Anchor | SURGE |
| One response compares sets correctly. The others compare sorted characters, lowercase the strings, or compare only set sizes. | Two responses count character multiplicities; two compare sets after lowercasing. These reject valid repetitions or erase case distinctions. | Three responses compare sets directly, ignoring repetition while preserving case. One response still sorts lowercased strings. |
| Observed correct: 1/4 | Observed correct: 0/4 | Observed correct: 3/4 |
| Passing implementation: return set(s0) == set(s1) | ||
| Method | Accuracy (%) | Reasoning tokens |
|---|---|---|
| Anchor | 50.83 | 8,065 |
| Donor | 47.50 | 7,651 |
| Best tested non-endpoint scalar ( ) | 49.17 | 9,141 |
| Linear interpolation ( ) | 47.50 | 7,659 |
| Late checkpoint average (3 checkpoints) | 50.83 | 8,078 |
| Anchor-only spectral truncation | 44.17 | 8,320 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Setting | Value |
| RL | Algorithm / framework | GRPO / veRL |
| Training data | DAPO-Math-17k | |
| Reward | Binary correctness, DAPO rule-based verifier | |
| Learning rate | , constant | |
| Train batch / PPO mini-batch | 256 / 64 | |
| Micro-batch per GPU | 1 |
| Variant | Accuracy (%) | Reasoning tokens | Cap (%) | |
|---|---|---|---|---|
| Anchor-only truncation | 0.329 | 44.17 | 8,320 | 2.5 |
| Scalar | 0.500 | 49.17 | 9,141 | 3.3 |
| Scalar | 0.250 | 47.50 | 7,723 | 0.0 |
| Scalar | 0.250 | 45.00 | 8,521 | 1.7 |
| Scalar | 0.500 | 44.17 | 8,028 | 0.8 |
| Scalar | 0.710 | 47.50 | 7,659 | 1.7 |
| Benchmark | Delta-SVD | Absolute-weight SVD |
|---|---|---|
| AIME24 | 54.17 / 7,798 | 41.67 / 8,173 |
| AIME25 | 38.33 / 7,383 | 35.00 / 7,368 |
| HMMT25 | 23.33 / 7,693 | 19.17 / 7,331 |
| BRUMO25 | 51.67 / 7,285 | 50.83 / 6,941 |
| CMIMC25 | 25.63 / 8,700 | 23.75 / 8,415 |
| AMC23 | 92.50 / 5,225 | 88.75 / 4,947 |
| Construction | Accuracy (%) | Reasoning tokens |
|---|---|---|
| 1 | 50.83 | 7,239 |
| 2 | 45.83 | 8,182 |
| 3 | 47.50 | 7,205 |
| 4 | 45.00 | 7,668 |
| Mean | 47.29 | 7,574 |
| Benchmark | SURGE | Anchor | Donor |
|---|---|---|---|
| AIME24 | / 40.97 | / 63.19 | / 35.42 |
| AIME25 | / 8.33 | / 22.22 | / 7.64 |
| HMMT25 | / 33.33 | / 11.11 | / 2.08 |
| BRUMO25 | / 2.78 | / 2.08 | / 11.11 |
| CMIMC25 | / 4.30 | / 18.75 | / 7.42 |
| AMC23 | / 9.38 | / 6.25 | / 13.67 |
| Family | Cells | Runs | Pooled |
|---|---|---|---|
| DeepSeek | 27 | 108 | |
| Nemotron | 27 | 108 | |
| Both | 54 | 216 |
| Benchmark | DeepSeek | Nemotron | Pooled |
|---|---|---|---|
| AIME24 | |||
| AIME25 | |||
| HMMT25 | |||
| BRUMO25 | |||
| CMIMC25 | |||
| AMC23 |
| DeepSeek | Nemotron | |||
|---|---|---|---|---|
| FC / reference | Accuracy (%) | Tokens | Accuracy (%) | Tokens |
| 0.0 (donor) | 47.50 | 7,651 | 60.00 | 9,212 |
| 0.1 | 47.50 | 7,517 | 62.50 | 10,060 |
| 0.2 | 48.33 | 7,639 | 65.00 | 10,420 |
| 0.3 | 49.17 | 7,439 | 66.67 | 10,786 |
| 0.4 | 48.33 | 8,096 | 66.67 | 10,841 |
| FC | Retained energy | Displacement | Accuracy (%) | Tokens |
|---|---|---|---|---|
| 0.1 | 0.509 | 0.929 | 81.7 | 4,258 |
| 0.2 | 0.655 | 0.886 | 82.5 | 4,146 |
| 0.3 | 0.754 | 0.846 | 81.6 | 4,338 |
| 0.4 | 0.828 | 0.807 | 83.4 | 4,462 |
| 0.5 | 0.883 | 0.766 | 82.9 | 4,498 |
| 0.6 | 0.924 | 0.724 | 81.1 | 4,591 |
| Problem. Mr. Potato Head has three hairstyles (or can be bald), two eyebrow pairs, one eye pair, two ear pairs, two lip pairs, and two alternative shoe pairs. All facial parts and one complete shoe pair are required. How many appearances are possible? Ground truth: . | |
| Donor | Anchor |
| Treats mandatory paired parts as optional left/right choices and uses an unsupported product of options. The resulting count does not respect the specified complete pairs. | Recognizes baldness but inconsistently counts the remaining features and shoe alternatives. Its final product undercounts the valid combinations. |
| Answer: (incorrect) | Answer: (incorrect) |
| Selected response: 2,677 reasoning tokens | Selected response: 3,013 reasoning tokens |
| Observed correct: 0/4 | Observed correct: 0/4 |
| SURGE, FC = 0.8 | |
| Problem. Insert parentheses in to minimize its value, preserving every number, operator, and their order. Ground truth: . | |
| Donor | Anchor |
| Introduces a subtraction inside a group containing 5 and 6, changing the original plus sign. Its candidate therefore solves a different expression. | Stops at the example value 3 after considering restricted groupings, while claiming that all possibilities have been checked. It fails to place the entire suffix under the minus sign. |
| Answer: (incorrect) | Answer: (incorrect) |
| Selected response: 6,988 reasoning tokens | Selected response: 10,339 reasoning tokens |
| Observed correct: 1/4 | Observed correct: 1/4 |
| SURGE, FC = 0.8 | |
| Problem. A square has diagonal 1. Points and satisfy , and . Find the distance between the incenters of and . Ground truth: . | |
| Donor | Anchor |
| Misreads the angle condition as involving the square diagonal, declares a conflict with its angle, and then appeals to unsupported symmetry to identify the two incenters. | Uses coordinates and an incenter calculation, but errors in the subsequent coordinate arithmetic lead to an incorrect distance. |
| Answer: (incorrect) | Answer: (incorrect) |
| Selected response: 10,119 reasoning tokens | Selected response: 17,122 reasoning tokens |
| Observed correct: 0/4 | Observed correct: 0/4 |
| SURGE, FC = 0.8 | |
| Problem. Find the largest real solution of . Ground truth: . | |
| Donor | Anchor |
| Tests isolated candidates and extrapolates an unsupported arithmetic pattern. Failure of a later candidate is treated as evidence that the preceding one is maximal. | Sets up the piecewise parametrization, but decimal-division errors underestimate the admissible integer ranges. It stops at a valid, nonmaximal solution. |
| Answer: (incorrect) | Answer: (incorrect) |
| Selected response: 14,224 reasoning tokens | Selected response: 17,295 reasoning tokens |
| Observed correct: 0/4 | Observed correct: 0/4 |
| SURGE, FC = 0.8 | |
| Task. Implement count_list(lst) to count top-level elements that are lists. For mixed inputs, counting all elements is incorrect. | |
| Donor | Anchor |
| Three responses return the outer length. One uses the correct type test but submits the wrong function name, count_lists , and fails the required interface. | Three responses return the outer length, fitting the all-list examples but failing mixed-type inputs. Only one tests whether each element is a list. |
| Observed correct: 0/4 | Observed correct: 1/4 |
| SURGE | |
| Three responses use isinstance(item, list) under the required function name. They count list-valued elements while ignoring scalars. One response still uses the outer-length shortcut. | |
| return sum(1 for item in lst if isinstance(item, list)) | |
| FC | 0.1 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek | 0.518 | 0.638 | 0.724 | 0.793 | 0.848 | 0.893 | 0.929 | 0.959 | 0.983 |
| Nemotron | 0.405 | 0.546 | 0.652 | 0.736 | 0.805 | 0.862 | 0.909 | 0.947 | 0.978 |
| OLMo | 0.509 | 0.655 | 0.754 | 0.828 | 0.883 | 0.924 | 0.954 | 0.976 | 0.991 |
| FC | FC | |||
|---|---|---|---|---|
| Anchor | Actual | Reference | Actual | Reference |
| DeepSeek | .929 | .854 | .959 | .913 |
| Nemotron | .909 | .854 | .947 | .913 |
| OLMo | .954 | .920 | .976 | .958 |
| FC | Absolute-weight SVD | Delta-SVD |
|---|---|---|
| 0.2 | 7.0 | 63.8 |
| 0.8 | 28.5 | 95.9 |
| Stage | DeepSeek | Nemotron |
| Read three checkpoints (warm page cache) | 0.04 | 0.03 |
| Transfer to GPU | 1.38 | 1.40 |
| 196 FP32 SVDs | 29.45 | 29.56 |
| Projection and fusion | 1.81 | 1.79 |
| Retained-energy curve | 0.07 | 0.07 |
| Displacement curve | 0.34 | 0.34 |
| Component / reconstruction | DeepSeek | Nemotron |
| Rollout, 32-way synchronous accounting | 1.60 | 2.46 |
| Rollout, larger-batch throughput accounting | 0.76 | 1.76 |
| Old-policy log probabilities | 0.14 | 0.32 |
| Actor forward/backward and four AdamW updates | 0.57 | 1.24 |
| Total, synchronous | 2.32 | 4.01 |
| Total, throughput-based | 1.48 | 3.31 |
| RL step | Role | Accuracy (%) | SD | Reasoning tokens |
|---|---|---|---|---|
| 0 | Initialization | 42.1 | 3.3 | 2,293 |
| 100 | 65.1 | 2.3 | 5,090 | |
| 200 | 67.2 | 4.1 | 4,521 | |
| 300 | 70.6 | 3.7 | 4,597 | |
| 400 | 75.0 | 3.6 | 4,609 | |
| 500 | 74.7 | 1.5 | 4,317 |