cs.LGOct 1, 2026

Adapter Thickets: Splitting an RLVR Budget Beats Concentrating It

Authors: Jonathan Williams, Esin Tureci Karthik R. Narasimhan

Organizations: Princeton University

Abstract

Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse for making each completion better. The standard pipeline composes the two: train one policy with RLVR, then sample it many times and vote. We show that this composition is lossy. A vote can only overturn mistakes that its voters do not share, and RLVR sharpens a policy so that its samples increasingly make the same mistakes. With every method drawing exactly 160160 completions per problem, training a single LoRA adapter on the full RLVR budget raises single-sample accuracy on every model we test (1.51.5B-88B). Yet on three of four models it leaves the majority vote below that of the untrained base model, by up to 4.84.8 points. The damage builds during training: voter errors grow steadily more correlated, and the majority vote accuracy peaks early before falling by up to 7.07.0 points. The cause is concentration, not RLVR itself. We split the same data and training budget across KK LoRA adapters, each trained on its own random disjoint shard, and call the result an adapter thicket. Thickets out-vote the fully trained adapter in all 1616 (model, KK) settings, and for K≥4K{\geq}4 they stay within 0.80.8 points of the base model or above it. A single adapter stopped early, at a thicket member's step count, is a strong control that matches thickets for small KK. For K≥8K{\geq}8, thickets keep more of RLVR's single-sample gain and out-vote this control in six of eight settings. The cost of concentration also grows with the number of votes: from 1616 to 160160 votes, the thicket's lead over the fully trained adapter widens from 1.31.3 to 3.33.3 points. When the plan is to sample and vote, an RLVR budget is better spent broad than deep.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

    Jul 28, 2026Pixel Nomand, Elena Voss, Marcus Hale +1Reinforcement Learning With Verifiable RewardToken Budget Allocation

  2. Diversifying RLVR Rollouts via First-Token Exploration

    May 27, 2026Soeun Kim, Albert NoReinforcement Learning With Verifiable RewardVerifiable Rewards

  3. When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

    May 19, 2026Yuchun Miao, Sen Zhang, Yuqi Zhang +4Reinforcement Learning With Verifiable RewardVerifiable Rewards