Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It
Organizations: Princeton University
Abstract
Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test (B-B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to points. The cause is concentration, not RLVR itself. We split the same data and training budget across LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all (model, ) settings, and for they stay within points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small . For , thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from to votes, the thicket's lead over the fully trained adapter widens from to points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.
Figures & tables
| Model | Method | ||||
|---|---|---|---|---|---|
| Qwen2.5-1.5B | Base | 51.3 0.4 | 51.3 0.4 | 51.3 0.4 | 51.3 0.4 |
| GRPO-single (matched) | 54.6 0.7 | 55.3 2.0 | 53.7 0.2 | 52.2 0.9 | |
| GRPO-single-adapter | 51.6 1.6 | 51.6 1.6 | 51.6 1.6 | 51.6 1.6 | |
| Thicket | 55.2 2.0 | 53.6 0.7 | 54.6 0.5 | 54.2 0.2 | |
| Qwen2.5-3B | Base | 63.8 0.3 | 63.8 0.3 | 63.8 0.3 | 63.8 0.3 |
| GRPO-single (matched) | 61.9 1.2 | 63.5 1.2 | 63.3 0.1 | 63.5 0.8 |
| Coverage, math | Coverage, code | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Method ( ) | 2 | 4 | 8 | 16 | 2 | 4 | 8 | 16 |
| Qwen2.5-1.5B | Base | 89.6 0.3 | 89.6 0.3 | 89.6 0.3 | 89.6 0.3 | 72.2 0.4 | 72.2 0.4 | 72.2 0.4 | 72.2 0.4 |
| GRPO-single (matched) | 89.6 0.4 | 89.3 0.5 | 90.0 0.3 | 89.5 0.5 | 69.0 1.8 | 70.2 0.9 | 72.0 0.5 | 71.9 0.3 | |
| GRPO-single-adapter | 87.7 0.7 | 87.7 0.7 | 87.7 0.7 | 87.7 0.7 | 65.8 1.4 | 65.8 1.4 | 65.8 1.4 | 65.8 1.4 | |
| Thicket | 88.8 1.0 | 89.5 0.4 | 89.6 0.5 | 89.8 0.6 | 69.2 1.4 | 70.6 0.6 | 70.6 0.5 | 70.3 0.3 | |
| Qwen2.5-3B | Base | 91.7 0.1 | 91.7 0.1 | 91.7 0.1 | 91.7 0.1 | 75.6 1.1 | 75.6 1.1 | 75.6 1.1 | 75.6 1.1 |
| Model | Shards | Majority vote | Filtered vote | Coverage | Voter acc. | |
|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | Subject | 54.2 | 55.2 | 89.6 | 34.4 | 0.37 |
| Random | 54.8 | 56.0 | 89.6 | 34.4 | 0.37 | |
| Qwen2.5-3B | Subject | 62.9 | 63.4 | 91.5 | 44.8 | 0.44 |
| Random | 63.1 | 62.7 | 91.0 | 44.9 | 0.44 | |
| Qwen2.5-7B | Subject | 69.1 | 69.3 | 92.7 | 55.2 | 0.51 |
| Random | 69.3 | 69.5 | 92.1 | 55.0 | 0.51 |
Appendix figures & tables31 assets
Supplementary material from the paper’s appendix.
Appendix
| LoRA rank / / dropout | 64 / 128 / 0.0 | prompts per step | 4 |
| LoRA targets | q,k,v,o,gate,up,down proj. | rollouts per prompt | 8 |
| learning rate | (constant, 5% warmup) | rollout temperature | 1.0 |
| weight decay / grad. clip | 0.01 / 0.1 | max prompt / completion | 1,024 / 2,048 tok. |
| KL coefficient | 0.0 | reward | binary verified |
| advantage | group-normalized (GRPO) | epochs | 1 |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 2 | 4 | 2 | 4 | 2 | 4 | 2 | 4 |
| Qwen2.5-1.5B | MATH500 | 67.7 0.3 | 67.7 0.3 | 68.5 1.3 | 69.9 1.0 | 67.5 1.7 | 67.5 1.7 | 68.6 1.0 | 69.1 0.6 |
| GSM8K | 86.4 0.4 | 86.4 0.4 | 84.1 0.3 | 85.1 0.3 | 82.4 1.1 | 82.4 1.1 | 84.6 0.1 | 85.1 0.4 | |
| OlympiadBench | 32.5 0.4 | 32.5 0.4 | 33.6 0.5 | 33.7 1.3 | 31.9 0.3 | 31.9 0.3 | 34.1 0.4 | 34.3 0.8 | |
| Countdown | 39.8 1.9 | 39.8 1.9 | 51.8 2.2 | 54.7 8.0 | 42.4 10.5 | 42.4 10.5 | 53.8 9.8 | 47.8 3.2 | |
| Knights & Knaves | 30.1 0.5 | 30.1 0.5 | 35.0 4.1 | 33.3 1.0 | 33.7 1.7 | 33.7 1.7 | 35.1 1.6 | 31.7 0.5 | |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 8 | 16 | 8 | 16 | 8 | 16 | 8 | 16 |
| Qwen2.5-1.5B | MATH500 | 67.7 0.3 | 67.7 0.3 | 68.7 0.2 | 67.8 0.2 | 67.5 1.7 | 67.5 1.7 | 69.0 0.9 | 68.8 1.0 |
| GSM8K | 86.4 0.4 | 86.4 0.4 | 86.0 0.2 | 86.2 0.3 | 82.4 1.1 | 82.4 1.1 | 85.6 0.2 | 85.9 0.1 | |
| OlympiadBench | 32.5 0.4 | 32.5 0.4 | 32.9 0.9 | 32.2 0.3 | 31.9 0.3 | 31.9 0.3 | 33.2 0.4 | 33.9 0.7 | |
| Countdown | 39.8 1.9 | 39.8 1.9 | 48.5 1.0 | 45.0 3.5 | 42.4 10.5 | 42.4 10.5 | 53.6 0.3 | 49.6 0.6 | |
| Knights & Knaves | 30.1 0.5 | 30.1 0.5 | 32.7 0.3 | 29.7 1.3 | 33.7 1.7 | 33.7 1.7 | 31.7 2.1 | 32.6 1.7 | |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 2 | 4 | 2 | 4 | 2 | 4 | 2 | 4 |
| Qwen2.5-1.5B | MATH500 | 69.7 0.5 | 69.5 0.6 | 69.6 1.0 | 71.4 1.0 | 68.1 1.3 | 67.1 0.7 | 70.9 0.4 | 71.9 0.8 |
| GSM8K | 86.4 0.3 | 86.4 0.2 | 84.1 0.3 | 85.0 0.3 | 82.4 1.1 | 82.4 1.2 | 85.0 0.4 | 85.0 0.2 | |
| OlympiadBench | 35.9 0.3 | 35.4 0.5 | 35.4 1.3 | 36.2 1.7 | 31.9 1.3 | 32.5 0.9 | 35.4 1.1 | 36.6 0.6 | |
| Countdown | 38.0 2.8 | 38.2 2.7 | 51.3 3.2 | 54.7 8.0 | 41.7 10.9 | 42.4 10.5 | 53.8 9.8 | 47.8 3.2 | |
| Knights & Knaves | 28.5 0.6 | 29.4 0.6 | 34.9 2.3 | 33.0 1.3 | 34.0 3.0 | 33.2 1.5 | 36.2 1.5 | 32.2 0.4 | |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 8 | 16 | 8 | 16 | 8 | 16 | 8 | 16 |
| Qwen2.5-1.5B | MATH500 | 69.9 0.6 | 69.5 0.4 | 70.9 0.2 | 70.8 0.7 | 67.3 1.0 | 67.2 0.7 | 71.0 0.4 | 70.0 0.7 |
| GSM8K | 86.1 0.2 | 86.3 0.4 | 85.8 0.2 | 86.3 0.2 | 82.4 1.1 | 82.4 1.1 | 85.3 0.5 | 86.3 0.6 | |
| OlympiadBench | 35.8 0.4 | 34.7 0.8 | 35.2 0.6 | 34.5 0.7 | 32.6 1.0 | 32.8 1.4 | 36.6 1.0 | 35.3 0.2 | |
| Countdown | 38.8 2.1 | 38.1 3.0 | 47.6 2.4 | 45.0 3.5 | 42.4 10.5 | 42.4 10.5 | 53.6 0.3 | 49.6 0.6 | |
| Knights & Knaves | 29.4 1.1 | 29.4 0.8 | 31.6 0.6 | 28.6 2.6 | 32.8 2.6 | 34.4 2.1 | 31.1 1.6 | 32.6 1.7 | |
| Model | MATH500 | GSM8K | OlympiadBench | Countdown | Knights & Knaves | Mean | |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | 2 | 70.9 0.4 | 85.0 0.4 | 35.4 1.1 | 53.8 9.8 | 36.2 1.5 | 56.3 1.7 |
| 4 | 71.9 0.8 | 85.0 0.2 | 36.6 0.6 | 47.8 3.2 | 32.2 0.4 | 54.7 0.9 | |
| 8 | 71.0 0.4 | 85.3 0.5 | 36.6 1.0 | 53.6 0.3 | 31.1 1.6 | 55.5 0.5 | |
| 16 | 70.0 0.7 | 86.3 0.6 | 35.3 0.2 | 49.6 0.6 | 32.6 1.7 | 54.8 0.2 | |
| Qwen2.5-3B | 2 | 77.9 0.3 | 91.1 0.6 | 45.3 1.4 | 47.6 6.1 | 53.7 1.3 | 63.1 1.7 |
| 4 | 78.7 1.1 | 91.5 0.2 | 45.3 0.6 | 52.6 2.6 | 55.3 3.7 | 64.7 0.4 |
| Model | MATH500 | GSM8K | OlympiadBench | Countdown | Knights & Knaves | Mean | |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | 2 | 69.7 0.5 | 86.4 0.3 | 35.9 0.3 | 38.0 2.8 | 28.5 0.6 | 51.7 0.6 |
| 4 | 69.5 0.6 | 86.4 0.2 | 35.4 0.5 | 38.2 2.7 | 29.4 0.6 | 51.8 0.4 | |
| 8 | 69.9 0.6 | 86.1 0.2 | 35.8 0.4 | 38.8 2.1 | 29.4 1.1 | 52.0 0.4 | |
| 16 | 69.5 0.4 | 86.3 0.4 | 34.7 0.8 | 38.1 3.0 | 29.4 0.8 | 51.6 0.3 | |
| Qwen2.5-3B | 2 | 78.9 0.4 | 91.6 0.5 | 45.7 1.3 | 51.7 1.5 | 54.2 3.4 | 64.4 0.4 |
| 4 | 78.8 0.4 | 91.6 0.2 | 45.6 0.9 | 51.0 1.5 | 54.4 3.0 | 64.3 0.1 |
| Model | MATH500 | GSM8K | OlympiadBench | Countdown | Knights & Knaves | Mean | |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | 2 | 68.1 1.3 | 82.4 1.1 | 31.9 1.3 | 41.7 10.9 | 34.0 3.0 | 51.6 1.7 |
| 4 | 67.1 0.7 | 82.4 1.2 | 32.5 0.9 | 42.4 10.5 | 33.2 1.5 | 51.5 1.8 | |
| 8 | 67.3 1.0 | 82.4 1.1 | 32.6 1.0 | 42.4 10.5 | 32.8 2.6 | 51.5 1.7 | |
| 16 | 67.2 0.7 | 82.4 1.1 | 32.8 1.4 | 42.4 10.5 | 34.4 2.1 | 51.8 1.6 | |
| Qwen2.5-3B | 2 | 74.4 1.5 | 90.0 0.3 | 42.3 1.9 | 35.7 18.1 | 53.2 5.0 | 59.1 5.3 |
| 4 | 74.3 2.0 | 89.8 0.5 | 42.3 2.4 | 35.7 18.1 | 52.5 5.2 | 58.9 5.5 |
| Model | MATH500 | GSM8K | OlympiadBench | Countdown | Knights & Knaves | Mean | |
|---|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | 2 | 69.6 1.0 | 84.1 0.3 | 35.4 1.3 | 51.3 3.2 | 34.9 2.3 | 55.0 0.0 |
| 4 | 71.4 1.0 | 85.0 0.3 | 36.2 1.7 | 54.7 8.0 | 33.0 1.3 | 56.1 2.0 | |
| 8 | 70.9 0.2 | 85.8 0.2 | 35.2 0.6 | 47.6 2.4 | 31.6 0.6 | 54.2 0.4 | |
| 16 | 70.8 0.7 | 86.3 0.2 | 34.5 0.7 | 45.0 3.5 | 28.6 2.6 | 53.1 0.7 | |
| Qwen2.5-3B | 2 | 77.2 0.3 | 90.9 0.3 | 44.3 0.7 | 46.2 4.5 | 54.6 2.8 | 62.6 1.3 |
| 4 | 76.6 0.7 | 91.1 0.6 | 44.4 1.4 | 50.8 3.5 | 56.0 2.2 | 63.8 1.6 |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 2 | 4 | 2 | 4 | 2 | 4 | 2 | 4 |
| Qwen2.5-1.5B | MATH500 | 93.5 0.8 | 93.5 0.8 | 94.0 0.7 | 93.9 0.3 | 92.7 1.1 | 92.7 1.1 | 92.8 0.3 | 93.7 0.2 |
| GSM8K | 99.1 0.1 | 99.1 0.1 | 98.6 0.2 | 99.0 0.2 | 98.1 0.3 | 98.1 0.3 | 98.9 0.3 | 98.9 0.1 | |
| OlympiadBench | 67.9 0.4 | 67.9 0.4 | 69.1 0.3 | 68.4 0.1 | 66.3 0.9 | 66.3 0.9 | 68.2 1.0 | 69.3 0.9 | |
| Countdown | 87.8 1.0 | 87.8 1.0 | 86.2 2.7 | 85.1 2.1 | 81.6 4.5 | 81.6 4.5 | 84.2 4.3 | 85.3 1.8 | |
| Knights & Knaves | 100.0 0.0 | 100.0 0.0 | 100.0 0.0 | 100.0 0.0 | 100.0 0.0 | 100.0 0.0 | 99.9 0.2 | 100.0 0.0 | |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 8 | 16 | 8 | 16 | 8 | 16 | 8 | 16 |
| Qwen2.5-1.5B | MATH500 | 93.5 0.8 | 93.5 0.8 | 94.1 0.4 | 92.5 0.4 | 92.7 1.1 | 92.7 1.1 | 93.9 0.5 | 93.4 0.2 |
| GSM8K | 99.1 0.1 | 99.1 0.1 | 99.1 0.2 | 99.1 0.1 | 98.1 0.3 | 98.1 0.3 | 98.8 0.2 | 98.8 0.2 | |
| OlympiadBench | 67.9 0.4 | 67.9 0.4 | 68.9 0.1 | 68.6 0.6 | 66.3 0.9 | 66.3 0.9 | 68.7 0.1 | 69.7 0.8 | |
| Countdown | 87.8 1.0 | 87.8 1.0 | 88.0 1.3 | 87.6 2.7 | 81.6 4.5 | 81.6 4.5 | 86.4 2.0 | 87.1 1.9 | |
| Knights & Knaves | 100.0 0.0 | 100.0 0.0 | 100.0 0.0 | 99.9 0.2 | 100.0 0.0 | 100.0 0.0 | 99.9 0.2 | 100.0 0.0 | |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 2 | 4 | 2 | 4 | 2 | 4 | 2 | 4 |
| Qwen2.5-1.5B | MBPP | 93.4 0.0 | 93.4 0.0 | 90.6 2.4 | 92.4 1.0 | 88.1 1.6 | 88.1 1.6 | 91.4 0.9 | 92.7 0.3 |
| HumanEval | 96.7 0.9 | 96.7 0.9 | 92.5 3.0 | 93.9 0.6 | 88.8 1.4 | 88.8 1.4 | 92.9 1.4 | 93.9 0.6 | |
| LeetCode | 26.5 0.3 | 26.5 0.3 | 24.0 3.7 | 24.4 2.1 | 20.5 2.6 | 20.5 2.6 | 23.4 2.8 | 25.1 1.1 | |
| Mean | 72.2 0.4 | 72.2 0.4 | 69.0 1.8 | 70.2 0.9 | 65.8 1.4 | 65.8 1.4 | 69.2 1.4 | 70.6 0.6 | |
| Qwen2.5-3B | MBPP | 95.2 0.5 | 95.2 0.5 | 93.0 2.8 | 92.9 1.8 | 90.8 1.1 | 90.8 1.1 | 92.6 1.3 | 93.1 1.8 |
| Base | GRPO-single (matched) | GRPO-single-adapter | Thicket | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Benchmark ( ) | 8 | 16 | 8 | 16 | 8 | 16 | 8 | 16 |
| Qwen2.5-1.5B | MBPP | 93.4 0.0 | 93.4 0.0 | 93.8 0.1 | 93.7 0.8 | 88.1 1.6 | 88.1 1.6 | 93.3 0.1 | 93.0 0.5 |
| HumanEval | 96.7 0.9 | 96.7 0.9 | 96.7 0.7 | 96.1 1.5 | 88.8 1.4 | 88.8 1.4 | 94.1 1.4 | 93.5 0.9 | |
| LeetCode | 26.5 0.3 | 26.5 0.3 | 25.6 0.9 | 25.9 0.4 | 20.5 2.6 | 20.5 2.6 | 24.3 0.3 | 24.6 1.5 | |
| Mean | 72.2 0.4 | 72.2 0.4 | 72.0 0.5 | 71.9 0.3 | 65.8 1.4 | 65.8 1.4 | 70.6 0.5 | 70.3 0.3 | |
| Qwen2.5-3B | MBPP | 95.2 0.5 | 95.2 0.5 | 94.8 0.5 | 95.6 0.2 | 90.8 1.1 | 90.8 1.1 | 94.2 0.9 | 95.4 0.7 |
| What full RLVR does | Thicket margin (MV, p.p.) | ||||
|---|---|---|---|---|---|
| Model | (pass@1) | thicket base | base GRPO-single-adapter | ||
| Qwen2.5-1.5B | +7.3 | +0.13 | +2.9 | 0.3 | +2.6 |
| Qwen2.5-3B | +1.0 | +0.07 | +0.3 | +4.8 | +5.1 |
| Qwen2.5-7B | +0.7 | +0.05 | 0.5 | +1.7 | +1.2 |
| Llama-3.1-8B | +9.7 | +0.19 | +1.4 | +4.1 | +5.5 |
| Voter accuracy (%) | Error correlation | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Suite | Model | Base | Matched | Single | Thicket | Base | Matched | Single | Thicket |
| Math | Qwen2.5-1.5B | 28.9 0.1 | 31.1 0.1 | 36.2 0.5 | 34.0 0.1 | 0.30 0.00 | 0.32 0.00 | 0.43 0.01 | 0.37 0.00 |
| Qwen2.5-3B | 44.1 0.1 | 44.1 0.2 | 45.1 0.8 | 44.8 0.1 | 0.42 0.00 | 0.42 0.00 | 0.49 0.01 | 0.43 0.00 | |
| Qwen2.5-7B | 55.2 0.1 | 55.2 0.0 | 55.9 0.4 | 55.1 0.0 | 0.51 0.00 | 0.51 0.00 | 0.56 0.01 | 0.51 0.00 | |
| Llama-3.1-8B | 29.2 0.0 | 30.0 0.2 | 38.9 0.2 | 35.3 0.0 | 0.27 0.00 | 0.28 0.00 | 0.46 0.01 | 0.34 0.00 | |
| Code | Qwen2.5-1.5B | 34.7 0.1 | 34.9 0.2 | 41.0 1.5 | 37.5 0.2 | 0.45 0.00 | 0.46 0.00 | 0.65 0.03 | 0.50 0.01 |
| Votes per problem (first responses of each method) | |||||||
| 16 | 32 | 48 | 64 | 96 | 128 | 160 | |
| Mean margin (p.p.) over the 16 (model, ) settings (settings with a positive margin) | |||||||
| GRPO-single-adapter Base | +0.15 (8) | 0.40 (8) | 1.08 (4) | 1.43 (4) | 1.93 (0) | 2.28 (4) | 2.58 (4) |
| Thicket GRPO-single-adapter | +1.33 (15) | +1.68 (16) | +2.01 (16) | +2.25 (16) | +2.86 (16) | +3.26 (16) | +3.28 (16) |
| Thicket Base | +1.48 (13) | +1.28 (13) | +0.93 (10) | +0.83 (8) | +0.94 (11) | +0.98 (10) | +0.71 (9) |
| Thicket GRPO-single (matched) | +0.56 (11) | +0.47 (11) | +0.28 (10) | +0.28 (9) | +0.23 (9) | +0.34 (11) | +0.27 (10) |
| Model | Method | ||||
|---|---|---|---|---|---|
| Qwen2.5-1.5B | Base | 44.6 0.3 | 44.6 0.3 | 44.6 0.3 | 44.6 0.3 |
| GRPO-single (matched) | 47.9 1.4 | 46.6 1.7 | 46.1 0.4 | 45.0 0.1 | |
| GRPO-single-adapter | 46.0 1.7 | 46.0 1.7 | 46.0 1.7 | 46.0 1.7 | |
| Thicket | 48.1 1.0 | 47.5 0.9 | 46.9 0.6 | 46.4 1.0 | |
| Qwen2.5-3B | Base | 56.9 0.9 | 56.9 0.9 | 56.9 0.9 | 56.9 0.9 |
| GRPO-single (matched) | 56.9 1.0 | 57.7 0.8 | 57.8 0.7 | 56.6 0.9 |
| Coverage, math | Coverage, code | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Method ( ) | 2 | 4 | 8 | 16 | 2 | 4 | 8 | 16 |
| Qwen2.5-1.5B | Base | 73.1 0.4 | 73.1 0.4 | 73.1 0.4 | 73.1 0.4 | 61.3 1.1 | 61.3 1.1 | 61.3 1.1 | 61.3 1.1 |
| GRPO-single (matched) | 74.7 0.8 | 74.4 1.5 | 73.7 0.7 | 73.3 0.5 | 61.4 2.2 | 61.4 1.1 | 62.5 0.8 | 61.9 0.9 | |
| GRPO-single-adapter | 72.8 1.5 | 72.8 1.5 | 72.8 1.5 | 72.8 1.5 | 58.2 0.8 | 58.2 0.8 | 58.2 0.8 | 58.2 0.8 | |
| Thicket | 74.5 0.6 | 74.2 0.8 | 74.3 0.6 | 73.9 0.7 | 61.4 1.9 | 62.1 1.1 | 62.4 0.7 | 63.5 0.3 | |
| Qwen2.5-3B | Base | 80.6 0.0 | 80.6 0.0 | 80.6 0.0 | 80.6 0.0 | 68.1 0.5 | 68.1 0.5 | 68.1 0.5 | 68.1 0.5 |
| Majority vote | Filtered vote | ||||
|---|---|---|---|---|---|
| Model | Benchmark | Subject | Random | Subject | Random |
| Qwen2.5-1.5B | MATH500 | 69.4 | 69.4 | 71.3 | 71.5 |
| GSM8K | 86.0 | 85.8 | 85.5 | 85.8 | |
| OlympiadBench | 32.6 | 33.6 | 35.4 | 37.8 | |
| Countdown | 49.3 | 49.8 | 49.3 | 49.8 | |
| Knights & Knaves | 33.8 | 35.2 | 34.4 | 35.2 | |