Measuring and Mitigating Solution Mode Collapse in RLVR
Organizations: Princeton University
Abstract
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.
Figures & tables
| pass@8 | PCMD | |||||||||||
| Graph | Count down | Python | MathIR | Pantry Plan | Avg | Graph | Count down | Python | MathIR | Pantry Plan | Avg | |
| Re:Dr (ours) | ||||||||||||
| Re:Max (ours) | ||||||||||||
| GAPO | ||||||||||||
| SetPO | ||||||||||||
| MaxRL (no replay) | ||||||||||||
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Generated answer / check | Solution mode | Equivalent responses | Certified modes |
|---|---|---|---|---|
| Graph coloring | Complete assignment; check fixed colors and edges | Color vector | Spacing only | 6.27 4–12 |
| Countdown | Parse operands and execute exactly | Normalized AST | Commutative order | 4.52 2–8 |
| Python factors | Execute restricted lambda externally | Return vector | Programs with the same return vector | 229.44 16–1,512 |
| MathIR | Execute bounded action program | State trajectory | Programs with the same states | 5.00 5 |
| PantryPlan | Check ingredient allocation and all constraints | Ingredient support | Quantity and ordering aliases | 18.19 8–40 |
| Domain | Level 1 | Level 2 |
|---|---|---|
| Graph coloring | Three hidden vertices | Four hidden vertices |
| Countdown | Three operands | Four operands, with values up to 14 |
| Python factors | Four executed cases, all at most 96 | Four cases, including a value above 96 |
| MathIR | One-sided linear families | Variable-on-both-sides families |
| PantryPlan | Base dietary and composition constraints | More active exclusions and tighter composition constraints |
| Component | Setting |
|---|---|
| Data | 384 training and 128 disjoint evaluation prompts per domain and level; evaluation prompts are never used for training or replay. |
| Fresh rollouts | , temperature , top- , with one prompt group per optimizer update. |
| Training horizon | Eight passes over the training set, totaling 3,072 optimizer updates, with one PPO epoch per batch. |
| Evaluation | Four groups of eight responses per prompt at temperature and top- at passes ; final comparisons use the last checkpoint. |
| Replay | At most solution modes per prompt, one stored exemplar per mode, deterministic round-robin buffer selection, and . |
| Training seeds | Five seeds per model–domain condition, with paired comparisons using common seeds whenever available. |
| Model | Comparison | pass@8 | PCMD |
|---|---|---|---|
| Qwen2.5-0.5B | Re:Dr Dr.GRPO | ||
| Falcon3-1B | Re:Dr Dr.GRPO | ||
| Qwen2.5-3B | Re:Dr Dr.GRPO | ||
| Qwen2.5-0.5B | Re:Max MaxRL | ||
| Falcon3-1B | Re:Max MaxRL | ||
| Qwen2.5-3B | Re:Max MaxRL |
| pass@8 | PCMD | |||
|---|---|---|---|---|
| Domain | MaxRL | Re:Max | MaxRL | Re:Max |
| Countdown | .631 | .805 | .001 (81) | .138 (104) |
| Graph | .396 | .818 | .055 (51) | .353 (104) |
| MathIR | .000 | .744 | – (0) | .017 (96) |
| PantryPlan | .760 | .779 | .384 (97) | .348 (101) |
| Domain | Control PCMD | Online RLEP PCMD | Re:Dr PCMD | PCMD | Replay share |
|---|---|---|---|---|---|
| Graph | .003 | .344 | .711 | +.340 | .71 |
| Countdown | .017 | .054 | .538 | +.037 | .66 |
| Python | .000 | .000 | .111 | .000 | .20 |
| MathIR | .000 | .013 | .018 | +.013 | .58 |
| PantryPlan | .000 | .254 | .484 | +.254 | .93 |
| Mean | .004 | .133 | .372 | +.129 | .62 |
| Estimate | Replay lower | No change | Replay higher |
|---|---|---|---|
| Pooled estimate | 27/29 | 1/29 | 1/29 |
| Disjoint split 0 | 25/29 | 3/29 | 1/29 |
| Disjoint split 1 | 24/29 | 3/29 | 2/29 |
| Domain | Mixed groups |
|---|---|
| Graph | .144 |
| Countdown | .140 |
| Python | .015 |
| MathIR | .364 |
| PantryPlan | .123 |
| Control subset | pass@8 | PCMD | |
|---|---|---|---|
| All comparisons | 15 | +.352 | +.189 |
| Increasing control pass@8 | 9 | +.379 | +.118 |
| Nonzero initial pass@8 , mixed groups | 6 | +.247 | +.136 |
| Domain | Selected setting | Selected pass@8 | Baseline pass@8 |
|---|---|---|---|
| Graph | LR | .429 | .325 |
| Countdown | entropy | .635 | .493 |
| Python | baseline | .586 | .586 |
| MathIR | entropy | .490 | .460 |
| PantryPlan | LR | .600 | .531 |
| Domain | Control PCMD | Re:Dr PCMD | PCMD | pass@8 |
|---|---|---|---|---|
| Graph | .054 | .764 | ||
| PantryPlan | .091 | .563 |
| Benchmark | Base pass@32 | Dr.GRPO | Re:Dr | Dr.GRPO avg@32 | Re:Dr avg@32 |
|---|---|---|---|---|---|
| MATH-500 | .894 | .889 | .909 | .470 | .537 |
| AMC23 | .875 | .767 | .850 | .261 | .339 |
| OlympiadBench | .628 | .546 | .620 | .196 | .234 |
| Minerva | .474 | .491 | .540 | .135 | .167 |
| AIME24 | .367 | .333 | .383 | .057 | .081 |
| AIME25 | .200 | .211 | .217 | .021 | .025 |
| Benchmark | Base | Dr.GRPO | MaxRL | Re:Max |
|---|---|---|---|---|
| MATH-500 | .894 | .889 | .902 | .904 |
| AMC23 | .875 | .767 | .850 | .850 |
| AIME24 | .367 | .333 | .333 | .300 |
| AIME25 | .200 | .211 | .167 | .200 |
| Model | Cell | pass@8 : benchmark neutral | PCMD : benchmark neutral |
|---|---|---|---|
| Dr.GRPO | Python L2 | – | |
| Dr.GRPO | Python L3 | – | |
| Re:Dr | Python L2 | ||
| Re:Dr | Python L3 | ||
| Dr.GRPO | MathIR L2 | ||
| Re:Dr | MathIR L2 |
| One withdrawal | Two withdrawals | |||
|---|---|---|---|---|
| Domain | Re:Dr | Re:Max | Re:Dr | Re:Max |
| Graph | 19.65 | 10.99 | 23.65 | 12.51 |
| Countdown | 9.58 | 0.10 | 6.11 | -4.87 |
| Python | 7.67 | 5.68 | 10.10 | 8.16 |
| MathIR | 0.14 | 0.09 | 0.20 | 0.12 |
| PantryPlan | 9.78 | 6.61 | 7.23 | 4.94 |
| Domain | saved | by 8 | calls |
|---|---|---|---|
| Graph | 67.58 | 68.36 | -5.37 |
| Countdown | 6.76 | 8.11 | -0.55 |
| Python | 55.86 | 63.67 | -4.67 |
| MathIR | 20.54 | 20.98 | -1.68 |
| PantryPlan | 16.41 | 20.70 | -1.44 |
| Level | (%) | (%) | |
|---|---|---|---|
| 1 | 95.8 | 1.71 | 74.0 [71.9, 76.1] |
| 2 | 77.5 | 1.16 | 82.3 [80.3, 84.4] |
| 3 | 80.3 | 1.18 | 84.5 [82.8, 86.2] |
| Macro PCMD by level | |||||
|---|---|---|---|---|---|
| Deployment | Macro PCMD | L1 | L2 | L3 | |
| DeepSeek V4 Pro | .320 | 1.47 | .289 | .319 | .232 |
| Kimi K3 | .302 | 1.43 | .374 | .307 | .253 |
| GPT-5.4 | .264 | 1.36 | .338 | .247 | .224 |
| Grok 4.3 | .239 | 1.31 | .275 | .179 | .171 |
| GPT-5.6 Sol | .226 | 1.29 | .256 | .229 | .183 |
| Level 1 | Level 2 | Level 3 | ||||
|---|---|---|---|---|---|---|
| Deployment | pass@8 | # modes | pass@8 | # modes | pass@8 | # modes |
| GPT-5.6 Sol | 97.5 | 1.90 | 98.2 | 1.59 | 99.4 | 1.54 |
| Claude Opus 5 | 97.9 | 1.42 | 99.6 | 1.45 | 99.8 | 1.35 |
| GPT-5.4 | 97.7 | 2.00 | 98.0 | 1.68 | 99.4 | 1.62 |
| Grok 4.3 | 99.4 | 2.12 | 100.0 | 1.70 | 100.0 | 1.68 |
| Kimi K3 | 98.0 | 2.26 | 100.0 | 1.95 | 100.0 | 1.79 |
| Deployment | Level | Refusals | Accuracy (%) | |
|---|---|---|---|---|
| Claude Opus 5 | 1 | |||
| Claude Opus 5 | 2 | |||
| Claude Opus 5 | 3 | |||
| GPT-5.6 Sol | 1 | |||
| GPT-5.6 Sol | 2 | |||
| GPT-5.6 Sol | 3 |
| Deployment | pass@8 (%) | # modes | PCMD |
|---|---|---|---|
| GPT-5.6 Sol | |||
| GPT-5.4 | |||
| Grok 4.3 | |||
| Kimi K3 | |||
| Claude Opus 4.8 |
| pass@8 (%) | distinct@8 | (%) | |
|---|---|---|---|
| 0.0 | 65.6 | .673 | 62.1 |
| 0.5 | 72.3 | .840 | 61.2 |
| 1.0 | 75.2 | .963 | 59.1 |
| 1.5 | 79.8 | 1.125 | 57.7 |
| 2.0 | 79.4 | 1.146 | 55.9 |
| Deployment | Accuracy (%) | distinct@8 |
|---|---|---|
| Grok 4.3 | ||
| Kimi K3 |