cs.LGOct 7, 2026

Decoupling Exploration from Optimization in RLVR

Authors: Saif Punjwani, Micah Goldblum

Organizations: Department of Computer Science, Columbia University · Department of Electrical Engineering, Columbia University

Abstract

Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@kk scaling, indicating that ExpDis produces models that generate more diverse correct solutions.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

    Jun 23, 2026Wenyang Hu, Junxiang Jia, Zhen Shu +3Verifiable RewardsTrajectory-Level Credit

  2. Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

    Oct 5, 2025Zishang Jiang, Jinyi Han, Tingyun Li +7Reinforcement Learning With Verifiable RewardVerifiable Rewards