cs.AISep 30, 2026

Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

Authors: Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman

Organizations: Department of Robotics and Mechatronics Engineering, University of Dhaka

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs

    May 27, 2026Yue Cheng, Jiajun Zhang, Xiaohui Gao +3Reinforcement Learning With Verifiable RewardLLM Reasoning Strategies

  2. On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

    Feb 16, 2026Yu Huang, Zixin Wen, Yuejie Chi +4Reinforcement Learning With Verifiable RewardCurriculum Reinforcement Learning