Reinforcement Learning with Verifiable Rewards

Also known as RLVR

Momentum

59 papers in the last four weeks, up 119% on the four weeks before. 0.6% of all new papers.

Jul 13Week of Sep 28

Latest papers 438

All topics
CardsList
  1. On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

    Feb 16, 2026Yu Huang, Zixin Wen, Yuejie Chi +4RL for Language Model ReasoningReinforcement Learning with Verifiable Rewards

  2. To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models

    Feb 13, 2026Haoqing Wang, Xiang Long, Ziheng Li +3RL for Language ModelsGradient Interference

  3. Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation

    Feb 10, 2026Pei-Chi Pan, Yingbin Liang, Sen LinReward ModelingReward Hacking

  4. F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare

    Feb 6, 2026Daniil Plyusov, Alexey Gorbatovski, Boris Shaposhnikov +4Reinforcement LearningRL for Language Model Reasoning

  5. Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation

    Feb 5, 2026Zhiqi Yu, Zhangquan Chen, Mengting Liu +2RL for Language Model ReasoningGroup Relative Policy Optimization

  6. CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

    Feb 4, 2026Xiao Zhu, Xinyu Zhou, Boyu Zhu +5RL for Code GenerationReward Modeling

  7. HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning

    Jan 30, 2026Weiqi Wang, Xin Liu, Binxuan Huang +13RL for Language ModelsRL for Language Model Reasoning

  8. Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

    Jan 16, 2026Lecheng Yan, Ruizhe Li, Guanhua Chen +5Reward HackingMemorization in Language Models

  9. ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

    Dec 2, 2025Salman Rahman, Sruthi Gorantla, Arpit Gupta +3Process Reward ModelsReinforcement Learning with Verifiable Rewards

  10. Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation

    Nov 14, 2025Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan LiRL for Language ModelsLLM Hallucination Mitigation

  11. Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs

    Oct 5, 2025Zishang Jiang, Jinyi Han, Tingyun Li +7RL for Language Model ReasoningLLM-Guided RL

  12. Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers

    Oct 1, 2025Xin-Qiang Cai, Wei Wang, Feng Liu +3Reinforcement LearningRL for Language Model Reasoning

  13. Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

    Aug 13, 2025Maxime Heuillet, Yufei Cui, Boxing Chen +2RL for Language Model ReasoningEfficient Language Model Training

  14. Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

    Date pendingMuhammad Khalifa, Zohaib Khan, Omer Tafveez +2RL BenchmarksReward Hacking